Introduction
Apache Cassandra is a popular NoSQL database that provides high availability and fault tolerance through its distributed architecture. However, many enterprises rely solely on Cassandra’s built-in replication mechanisms and fail to implement an effective backup strategy. While Cassandra’s replication can provide resiliency against node failures, it does not protect against data loss due to accidental deletions, application errors, or other disasters.
In this blog post, we will explore why Apache Cassandra should be backed up and how to create an effective Cassandra backup and recovery strategy.

Why should you backup an Apache Cassandra cluster?
1. Cassandra Data Loss Protection
Despite the high availability and fault tolerance offered by Apache Cassandra, data loss can still occur due to a variety of reasons such as human error, malicious attacks, or natural disasters. Backing up the Cassandra database regularly is essential to prevent any catastrophic data loss. With a proper Cassandra backup and recovery strategy in place, enterprises can quickly recover from data corruption or loss and minimize the impact on business operations.
2. Compliance Requirements
Many industries have strict compliance requirements for data retention and protection. For instance, healthcare organizations are required to maintain patient data for a certain period of time, while financial institutions must retain financial records for several years. In such cases, backing up Apache Cassandra regularly is mandatory to meet regulatory compliance requirements.
3. Disaster Recovery
Enterprises can suffer from a range of disasters such as fires, floods, or cyber-attacks that can result in the loss of data. In such scenarios, having a backup of Cassandra and a disaster recovery strategy can prove to be a lifesaver. By restoring the data from a backup, enterprises can quickly resume their operations and minimize the impact of the disaster on the business.
4. Granular Recovery
Backing up Apache Cassandra can also enable granular recovery of specific data sets or tables. This is particularly useful in situations where only a subset of the data has been corrupted or lost. With a backup, enterprises can restore only the required data instead of restoring the entire database.
5. Testing and Development
Backing up Apache Cassandra can also facilitate testing and development activities. By restoring a Cassandra backup, developers can create a copy of the production database in a testing environment, enabling them to test new features or application updates without affecting the live data. This can significantly reduce the risk of errors and downtime caused by faulty updates.
What type of failures can occur with Apache Cassandra?
Physical Failures
Even if you are running your workload in the public cloud, there are real physical systems behind the seemingly infinite and API-driven abstractions that engineers are now used to. These are some of the physical failure domains that you will need to cater for, in order to plan your DR strategy.
Server Failures
Bare metal servers and the servers behind your cloud instances or virtual machines will fail. You may be using the remote storage volumes for your Cassandra instances (generally slow and expensive), or local SSDs (ephemeral storage) with higher chances of data loss. Either way, you still need to implement your enterprise architecture with the assumption that these can disappear at any point due to failures.
Data Center Failures
Even a whole data center can have an outage. Some of the common ones I have seen are aircon failures that cause machines to shut down due to heat building up in the buildings. You may have your workload spread across three availability zones. However, they tend to be relatively near each other.
In 2022 one of the availability zones in London experienced an outage due to the unprecedentedly hot summer, and the air conditioning could not keep up (GCP London). Luckily the other two availability zones survived. However, I would suspect this was a pure stroke of luck since the cooling design would likely have been similar across all three zones and the weather conditions did not vary significantly between them.
Entire AWS / GCP regions are known to experience complete outages. The most recent high profile is the GCP Paris region becoming unavailable for a prolonged period.
Even if a problem is contained within one availability zone, the entire userbase attempt to migrate the workload to other AZs in the same region all at the same time, causing outages to the APIs, provisioning, capacity problems etc.
Of course, you should also consider data centers outages due to fire (OVH), flooding (Hurricane Sandy), and other severe weather events that could take out the data centres where you’re running your workload and storing your critical data.
Cassandra provides multi-DC deployment model out of the box. You can mitigate most physical failures using this amazing feature. However, your keyspaces may be configured to have replicas in specific DCs and not others.
Human Errors and Accidents
As CIOs and CTOs, you may have implemented a strong automation culture already, leveraging DevOps/SRE/Infrastructure-as-Code practices. However, even in the most technologically enabled organizations, I have seen time and time again some gremlins that can cause issues with your production data.
One financial services provider leveraging Apache Cassandra to store large volumes of commodity trading data phoned up for help one day. One of the engineers accidentally executed the test environment database cleanup script against their production servers. This script shut down the Cassandra servers and deleted all the SSTable database files.
Unfortunately, this company did not back up their production Cassandra cluster to recover!
Depending on the human errors and the damage they cause, it is extremely difficult to recover from a script that deletes data files from all Cassandra servers in one hit.
Application Issues
Applications writing to Cassandra may inadvertently delete or overwrite some critical data. You will want to recover data from a specific point in time in the past.
Worried? You Should Be!
As business or technology leaders of enterprises, data is one of the challenges that keep you up at night. Fortunately, implementing a Cassandra backup strategy is not difficult, and it should be done as part of any enterprise deployment of Cassandra.
How to Backup Your Cassandra Cluster
Your Cassandra backup policy needs to account for the loss you can tolerate and the time you have to recover. The following checks apply to the backup process rather than a specific restore topology.
-
Set recovery requirements. Agree a recovery point objective (RPO) and recovery time objective (RTO) with the application owners. Use those requirements to choose the backup interval and retention period as well as where the backups are stored.
-
Take snapshots on the required nodes.
nodetool snapshotnormally flushes memtables before creating hard links to SSTable files on the node where it runs. Plan coverage across the cluster because one node’s snapshot is not a complete cluster backup or a transactionally consistent snapshot across all replicas. -
Store the files outside the node’s failure domain. Copy the complete snapshot directories to independent storage and verify that the files can be read back. A snapshot on the same disk will be lost with that disk. Keep the node and table identities with each backup so files from different replicas remain distinguishable.
-
Archive commitlogs when point-in-time recovery is required. Configure archiving before the failure and monitor whether completed segments reach backup storage. Copying the live commitlog directory occasionally does not establish a complete archive. The commitlog archiving guide explains the configuration and recovery requirements.
-
Preserve the schema and configuration. Keep the snapshot’s
schema.cqland export the required keyspace DDL throughcqlshusingDESCRIBE KEYSPACE your_keyspace. Record the Cassandra version and replication settings alongside the backup.nodetool describeclusterreports cluster information and schema versions rather than the DDL needed to recreate a keyspace. -
Restore into an isolated test environment. Choose a restore method appropriate to the source and target topology. Validate representative application queries and restored records while measuring how long the recovery takes. Avoid connecting a test restore to the live cluster or allowing applications to write to it during validation.
-
Automate backup verification and retention. Alert on failed uploads and missing archives as well as failed snapshot jobs. Check that retention keeps the snapshot and all subsequent commitlog segments required for each recovery window before deleting older backups.
Apache Cassandra’s backup documentation describes snapshots and incremental SSTable backups. Native incremental backups preserve newly flushed SSTables and need a suitable snapshot baseline for recovery. They are different from the commitlog archive used for point-in-time recovery.
Remember, the exact steps and commands might vary depending on the version of Cassandra you are using and the specific backup strategy you have in mind. Always consult the official documentation and relevant resources for your Cassandra version to ensure accurate backup procedures.
Easily backup and restore your Cassandra cluster with AxonOps
AxonOps provides an enterprise-grade Cassandra backup and restore solution for your Cassandra clusters as part of the one-stop operations management platform. An enterprise can immediately implement and easily maintain an effective Cassandra backup and restore process through a highly intuitive UI while ensuring any compliance requirements are quickly addressed.