All articles

Restore Cassandra Backups to a Separate Cluster

Restoring a Cassandra backup to a separate cluster gives you somewhere to test recovery without disturbing production. You can also use the restored data to investigate a problem or rehearse an upgrade before making changes to the original cluster.

You’ll need AxonOps configured to back up the source cluster and a completed remote backup created by AxonOps before following this procedure. The orchestrator uses AxonOps backup records to locate the files and map them to the destination nodes.

Restore Scenarios

  • Recovery testing lets you restore a production backup into an isolated cluster and check that the application can read the recovered data while measuring how long the recovery takes.
  • Disaster recovery uses replacement hosts when the original Cassandra nodes are unavailable and restores each source node’s backup to its corresponding destination node.
  • Staging and troubleshooting gives an authorised team a separate copy of the backed-up data to reproduce an application problem without changing the production cluster.
  • Upgrade rehearsals start by restoring and validating the backup on the same Cassandra release before testing the planned upgrade on the recovered cluster.
  • Moving to new infrastructure restores the backup onto different hosts with new addresses and cluster or DC/rack names while retaining the same node count and token ownership. A production move also needs a cutover plan for writes made after the backup.

Every scenario uses the fixed-node-count restore described below and preserves the original token ownership without moving tokens between nodes.

The destination environment will often have different IP addresses and its own cluster name. Its datacentre and rack names may differ too. Changing the configuration files alone can leave Cassandra unable to start because the restored system tables still describe the source environment.

We’ll work through a three-node example using AxonOps to restore the backup and update those names before Cassandra starts. We’ll also look at what happens inside the SSTables and the replication settings that need attention after a datacentre rename.

The commands follow the AxonOps orchestrated restore procedure. This is a source-checked walkthrough with illustrative addresses rather than a measured recovery test. The SSTable Tools README currently describes the tool as under active development and not yet operator-ready for production. Rehearse the complete procedure in an isolated environment before relying on it for recovery.

The Source and Destination Clusters

Our example uses Cassandra 4.1 with GossipingPropertyFileSnitch and Murmur3Partitioner. The destination runs the same Cassandra release as the source so we can test the restore without introducing an upgrade at the same time.

Node count and tokens must stay unchanged

The destination must have exactly the same number of nodes as the source datacentre in the backup. Every source node must map to exactly one destination node that retains its original tokens.

No nodes can be added or removed and no token movements can take place as part of this restore.

SettingSourceDestination
Cluster nameproductionrestore-validation
Datacentredc1dc-restore
Number of nodes33
Node 1 and rack10.10.0.11 in rack-a10.20.0.11 in restore-a
Node 2 and rack10.10.0.12 in rack-b10.20.0.12 in restore-b
Node 3 and rack10.10.0.13 in rack-c10.20.0.13 in restore-c

We’re changing the rack labels while preserving the grouping of nodes into racks. Redistributing nodes between racks would require a different plan because replica placement would change.

The runner maps one source datacentre to one destination datacentre per mapping file. A multi-DC recovery needs planning for each datacentre and its replication settings rather than splitting this mapping across several destination DCs.

Keep the destination isolated from production. Its seeds and advertised addresses must belong to the destination environment. Restrict network access so restored nodes and test clients cannot contact the original cluster. A different cluster name doesn’t replace that isolation.

Prepare the Backup and Target Nodes

Start with a completed AxonOps remote backup covering every source node in the selected datacentre. Retain the schema and topology information needed to interpret it. The Cassandra backup and recovery guide covers backup storage and retention in more detail.

This procedure restores the selected snapshot backup. It doesn’t replay subsequent writes or establish an application-wide transaction boundary across node snapshots. Recovery to a later timestamp needs a separate commitlog and point-in-time restore procedure with the required archives available.

Check the following prerequisites before running the orchestrator.

  • AxonOps is managing backups on the source cluster and the selected remote backup has completed for every source node being restored.
  • The destination has the same node count as the source backup with a one-to-one node mapping that preserves all token ownership.
  • Cassandra is stopped on every destination node and its configured data directories exist and are empty.
  • The destination has enough disk space for the restored data and temporary work performed by SSTable Tools.
  • Each destination has the required Cassandra runtime and Java version alongside systemd and the AxonOps packages installed below. The axon-agent package includes axon-cassandra-restore while sstable-tools is installed separately.
  • The bastion has axon-restore-runner and SSH access to every destination through root or passwordless sudo.
  • The AxonOps API key is available as AXONOPS_API_KEY and the backup storage configuration is held in a protected file.
  • The destination nodes can read the remote backup storage and have the necessary credentials and encryption keys.
  • Restored credentials and application data remain protected by the destination’s access controls.

Install the Packages on Each Destination Node

Add the AxonOps repository to the operating system on every destination Cassandra node before installing axon-agent and sstable-tools. The following commands follow the AxonOps restore tool installation instructions. Use the block for your operating system and skip repository creation if the same AxonOps repository is already configured.

Debian and Ubuntu with APT

Install the repository prerequisites and import the AxonOps signing key into a dedicated keyring. The signed-by setting restricts this repository to that keyring.

sudo apt-get update
sudo apt-get install -y curl gnupg ca-certificates

curl -fsSL https://packages.axonops.com/apt/repo-signing-key.gpg \
  | sudo gpg --dearmor -o /usr/share/keyrings/axonops.gpg

sudo tee /etc/apt/sources.list.d/axonops-apt.list <<'EOF'
deb [signed-by=/usr/share/keyrings/axonops.gpg] https://packages.axonops.com/apt axonops-apt main
EOF

sudo apt-get update
sudo apt-get install axon-agent sstable-tools

RHEL-Compatible Linux with YUM

Create the AxonOps repository file before installing the packages. The published YUM configuration below disables package and repository signature verification through gpgcheck=0 and repo_gpgcheck=0. Confirm that these settings meet your organisation’s package security requirements before using them.

sudo tee /etc/yum.repos.d/axonops-yum.repo <<'EOF'
[axonops-yum]
name=axonops-yum
baseurl=https://packages.axonops.com/yum/
enabled=1
repo_gpgcheck=0
gpgcheck=0
EOF

sudo yum install axon-agent sstable-tools

Installing axon-agent places the restore command at /usr/share/axonops/axon-cassandra-restore without requiring a separate restore package. The sstable-tools package installs /usr/bin/sstable-tools. Install axon-restore-runner separately on the bastion as described in the orchestration host setup before running the mapping commands later in this guide.

Configure the Destination Nodes

The runner checks for empty data directories but doesn’t stop Cassandra or erase existing data. Verify the destination hosts before preparing them and don’t clear a directory just to make a preflight check pass.

Set the cluster name and snitch in the existing cassandra.yaml on every destination node.

cluster_name: 'restore-validation'
endpoint_snitch: GossipingPropertyFileSnitch

The corresponding cassandra-rackdc.properties on node 1 contains the following values.

dc=dc-restore
rack=restore-a

Use restore-b on node 2 and restore-c on node 3. These are configuration excerpts rather than complete files. Review the destination’s seeds and network settings alongside its TLS certificates before continuing.

Map the Backup to the New Nodes

Run the following command from the bastion using your AxonOps organisation name. The interactive selection asks for the source cluster and datacentre followed by the backup and destination cluster name.

axon-restore-runner generate-mapping \
  --axonops-url https://dash.axonops.cloud \
  --org my-organization \
  --output restore-mapping.json

Select production and dc1 as the source for this example and restore-validation as the destination cluster name. A self-hosted AxonOps installation uses its own server URL.

The generated file contains the backup identity and one entry per source node. Review every entry before running a restore because destination addresses initially contain the source addresses. Change only the destination fields and keep the source agent IDs generated from the backup.

Here is an illustrative entry for node 1 inside the mapping’s nodes array. The UUID is an example and must be replaced by the actual generated source agent ID.

{
  "source_agent_id": "11111111-1111-4111-8111-111111111111",
  "source_ip": "10.10.0.11",
  "source_name": "cass-1",
  "source_datacenter": "dc1",
  "source_rack": "rack-a",
  "destination_ip": "10.20.0.11",
  "destination_datacenter": "dc-restore",
  "destination_rack": "restore-a"
}

Apply the remaining two rows from our topology table to nodes 2 and 3. The top-level destination.cluster must be restore-validation and agree with every target’s cassandra.yaml.

Validate the edited mapping locally before making any connections to the destination nodes.

axon-restore-runner validate-mapping restore-mapping.json

This catches invalid fields and duplicate destinations. It doesn’t prove that an address belongs to the intended environment or that the selected backup contains the data you need.

Check and Run the Restore

Run a read-only preflight using the storage configuration and SSH key prepared for this restore. The paths below are examples that need to match your installation.

axon-restore-runner restore restore-mapping.json \
  --storage-config-file /secure/path/storage-config.json \
  --ssh-user admin \
  --ssh-args='-i /secure/path/restore-key' \
  --cassandra-yaml /etc/cassandra/cassandra.yaml \
  --cassandra-lib-dir /usr/share/cassandra/lib \
  --dry-run

The YAML and library paths refer to files on each destination node. The runner reads the configuration separately on every host and checks that the mapping agrees with the cluster name and the DC/rack values in cassandra-rackdc.properties.

Review the complete plan and resolve failures before proceeding. A successful dry run checks the prerequisites without downloading the backup or updating agent identities. It can’t establish that the restored application data will pass your checks.

Run the same command without --dry-run when the plan is correct. The confirmation requires typing RESTORE before the runner submits a systemd job to each destination node. Each node downloads its own backup files directly from remote storage so the bastion doesn’t have to relay all the data.

Losing the bastion session doesn’t terminate a node’s restore once systemd has accepted its job. Reconnect using the local run record to check progress and read the recent logs.

axon-restore-runner status restore-mapping.run.json \
  --watch \
  --log-lines 10

Investigate the logs of any failed job and fix the cause before repeating the original restore command with --retry-failed. Matching jobs that are already running or have succeeded aren’t submitted again. Keep the run record and node logs with the recovery test results.

How SSTable Tools Changes the Names

The restored system.local table contains the source node’s cluster name and topology. Cassandra checks that stored information against its configuration during startup. Updating the destination YAML and rack properties while leaving conflicting values in system.local can therefore prevent the node from starting.

The runner uses AxonOps SSTable Tools to update the relevant cells while the destination Cassandra service is stopped. SSTable Tools is Apache-2.0 licensed and builds on the original SSTable Tools project. We’re grateful to the contributors whose work makes these tools available to the Cassandra community.

system.local columnRestored value on node 1Destination valueConfiguration source
cluster_nameproductionrestore-validationTarget cassandra.yaml
data_centerdc1dc-restoreValidated destination mapping
rackrack-arestore-aValidated destination mapping

The runner combines the changed columns in one CQL UPDATE which SSTable Tools applies in a private workspace using the matching Cassandra runtime. It writes a new SSTable containing the changes and leaves the original SSTable files unchanged.

Files read togetherCluster name visible to Cassandra
Original restored SSTables onlyproduction
Original SSTables plus the new SSTable with a newer cluster_name cellrestore-validation

Cassandra reconciles those cells by timestamp when reading the table. The new write needs an explicit timestamp in microseconds that is greater than the source maximum. Future-dated source cells can make the current clock unsuitable and require investigation before retrying the update.

Every SSTable contributing to system.local must be included. Cassandra can spread one table across several data directories and different SSTables can contain different cells of the same row. The runner selects the table across those directories and reads it back with the new SSTable included before marking the rewrite complete.

SSTable Tools uses release-specific adapters and requires the matching schema. Its documented system-table workflow requires version 1.0.5 or newer and currently supports Murmur3Partitioner. It follows the selected Cassandra installation’s output format settings rather than converting the data to an arbitrary format. Check the system-table instructions against the installed release before using it independently.

This example leaves the rewrite to the orchestrator. Publishing SSTables into a running node’s live directories can race with Cassandra’s own file management and mustn’t become a shortcut for changing a live cluster’s identity.

Replication and Authentication After a DC Rename

Changing system.local.data_center doesn’t update the DC names in a keyspace’s NetworkTopologyStrategy replication map. A restored keyspace might still request three replicas in dc1 even though the destination nodes now advertise dc-restore.

Plan that schema change before starting the restored cluster and include replicated system keyspaces such as system_auth where applicable. Establish an authorised recovery access procedure beforehand because authentication can fail when its data has no replicas in the newly named DC. Disabling authentication to work around the problem could expose the restored data.

Inspect the destination’s replication settings once it is running and you have the planned administrative access.

SELECT keyspace_name, replication
FROM system_schema.keyspaces;

The following setting gives an application keyspace named shop three replicas in our three-node destination.

ALTER KEYSPACE shop WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'dc-restore': 3
};

Run that change only on the isolated destination and adapt it to the actual replication requirements. ALTER KEYSPACE doesn’t copy missing replicas or prove that the restored data is complete. Review ownership and complete any required repair before allowing application traffic.

The application’s driver also needs dc-restore as its local datacentre together with destination contact points and credentials. A client still configured for dc1 can fail to find appropriate hosts even when the restored nodes are healthy.

Verify the Restored Cluster

The runner leaves Cassandra stopped so you can review the restore results before startup. Start the destination nodes using the recovery procedure prepared for this Cassandra version after every node’s restore and metadata update has succeeded.

Check the identity directly on each node rather than querying a single contact point and assuming it describes the whole cluster.

SELECT cluster_name, data_center, rack, host_id
FROM system.local;

Use nodetool status to check the expected nodes and rack membership. Compare the restored tokens and schema with the source backup records and investigate startup errors before connecting test applications.

CheckExpected result
Cluster name on all three nodesrestore-validation
Datacentre on all three nodesdc-restore
Rack labelsrestore-a, restore-b and restore-c on the intended nodes
Replication mapsDestination DC names with the intended replication factors
Known application partitionsExpected values from the selected recovery point
Client connectionsDestination contact points and local DC with working TLS and authentication
Production isolationNo destination seed or application connection reaches the source cluster

Query known partitions and exercise representative application reads before treating the recovery as successful. Include TTL behaviour in those checks because restoring an older SSTable doesn’t restart a value’s TTL. Data that has expired by the time of the restore can remain expired.

The runner also resets the destination’s AxonOps agent identity files when the cluster name changes so it can register separately. The agent has its own identity which is separate from Cassandra’s system.local.host_id. Check that the restored environment appears under the intended name in AxonOps before enabling its operational schedules.

Record the elapsed time through to successful application checks alongside the backup’s recovery point. That gives you a recovery test you can repeat and compare as the dataset grows. AxonOps backup and restore provides the backup history and orchestration used here while the isolated cluster gives you somewhere to check that the recovered data actually serves your application.

All articles