All articles

Apache Cassandra 6.0 Part 10 - Upgrade and Production Validation

Upgrade and Production Validation

When the first node restarts on Cassandra 6.0 you’ll want to compare more than its version number with the nodes still running the previous release. Its existing SSTables and schema need to work with the new runtime while application drivers continue sending traffic. The upgrade plan also needs to cover the monitoring and recovery tools you would use if that node develops a problem.

This series follows pre-release Cassandra 6.0 and the checks in this post reflect 6.0-alpha3. Use the release notes and compatibility guidance for the exact version being tested before turning this into an upgrade procedure. A later alpha or beta may change the requirements and a release candidate or GA version needs its own review when available. The branch changelog helps explain the development work but doesn’t provide a production release guarantee.

Keep the Cassandra upgrade separate from adopting optional features so you can identify which change affected the workload. Accord changes application coordination while compression and cursor compaction change different parts of the storage path. Direct I/O and repair policy need separate tests too alongside SAI and guardrails or a move to JDK 21 and Generational ZGC.

Change categoryInclude in the binary upgradeValidate as a separate controlled change
Cassandra binary and supported configurationYes, after compatibility and rolling-upgrade tests.N/A
Drivers and application behaviourInventory and test before any production node is upgraded.Upgrade individual drivers and services through their own release process.
Storage format and compatibility modeFollow the exact release documentation and do not advance the format early.Advance only after every node and operational tool has been validated.
Cursor compaction and direct I/OLeave at the known baseline unless the feature test already supports a change.Test table eligibility, disk path, page-cache impact, and rollback separately.
Accord, compression dictionaries, SAI, constraintsDo not make application semantic changes during the database rollout.Test each feature with a specific workload and application migration.
Automated repair, guardrails, JDK 21Retain the known operating model through the binary rollout.Introduce scheduler policy, reject thresholds, or JVM choice after baseline stability is established.

Build the upgrade inventory

Start by recording what the cluster is running so you have something concrete to compare with a node after its upgrade. Include the table definitions and storage characteristics in that inventory because old files and schema history can affect how the new version behaves.

Inventory areaRecord before testingWhy it changes the upgrade outcome
Cassandra and JavaNode version, exact package or image, JDK vendor and patch level, JVM options, heap, direct-memory configuration, and process limits.A JVM or package change can look like a Cassandra regression when the binary is not the source of the difference.
Topology and networkingData centers, racks, token allocation, replication, seed and discovery configuration, TLS, internode ports, and node replacement procedure.Bootstrap, streaming, repair, and metadata changes need the actual topology rather than an idealized ring diagram.
Schema and storageKeyspaces, table schemas, counters, collections, indexes, compaction strategies, compression, TTLs, table size, SSTable format, disk layout, and largest partitions.Storage compatibility, cursor eligibility, index cost, repair volume, and compaction behaviour depend on the real tables.
Application clientsDriver type and version, authentication, consistency levels, retry policies, prepared-statement use, request sizes, paging, timeouts, and service ownership.Client compatibility and retry behaviour can turn a short node restart into an application incident.
Operations toolingMetrics, logs, tracing, JMX, agents, backup, restore, repair controller, alerting, deployment automation, incident runbooks, and access controls.An upgrade is not safe if operators lose the ability to diagnose or recover the cluster.
Capacity baselineDisk free space by data directory, compaction backlog, repair freshness, read/write latency, timeout rate, CPU, memory, network, and workload seasonality.A node under existing pressure is a poor canary and can conceal the reason an upgrade test fails.

Store the inventory with the test results so the first upgraded node can be compared with the same recorded configuration and workload. Keep using it as you move through a rack and then the full cluster instead of reconstructing the original state from separate discussions.

Before the rolling upgrade

  • Confirm the supported source version and exact Cassandra 6.0 release with its required rolling order and release-specific instructions.
  • Test representative data with old SSTables and tombstones as well as indexes and repair history. Run client traffic while compaction and streaming are active and include the normal backup workload.
  • Check the Cassandra driver versions used by applications and maintenance scripts including infrequent jobs so older clients are represented in the test. Use the driver-version guardrail to check reported versions before production where appropriate.
  • Restore a representative backup with the tools you normally use and verify that the restored data can be read.
  • Test bootstrap and replacement alongside decommission and restart in the supported test stages. Include streaming and repair with compaction and schema changes as well as recovery from a node failure.
  • Confirm monitoring and logs work with Cassandra 6.0 together with JMX and profiling. Check backups and alerts alongside access control and the incident tools used by the team.
  • Record the ring and schema with configuration and disk space before upgrading. Include repair age and compaction state alongside table reads and writes with SSTables per read and tombstones scanned. Retain latency and timeout measurements with CPU and heap use as well as direct memory and network activity.
  • Identify and disable automation that can change schema or bootstrap and decommission nodes during the CMS transition. Include node moves and replacements as well as assassinate operations in that restriction.
  • Choose CMS members across the intended failure domains before starting so the initial one-member CMS can be expanded to at least three after initialization.

Rolling upgrade procedure

  1. Use the documented procedure and supported rolling order for the exact Cassandra 6.0 release you intend to install.
  2. Upgrade Cassandra without changing the JDK where the existing runtime is already supported. Cassandra 6.0 supports JDK 11 and 17 alongside 21 so a 4.0 or 4.1 cluster on JDK 8 needs a planned JDK move before or within its supported upgrade sequence. Otherwise keep drivers and storage compatibility settings unchanged along with repair policy and optional features.
  3. Start with one representative node that isn’t already affected by an incident or under disk pressure. Check compaction and repair load before selecting it so existing problems don’t obscure the upgrade comparison.
  4. Return that node to service and check client traffic with topology and schema and metadata state. Review logs and streaming alongside compaction and repair while verifying monitoring and backup operation and host resource use.
  5. Observe the node under normal traffic and permitted maintenance work before comparing its behaviour with nodes that haven’t yet been upgraded.
  6. Upgrade the next small group only after the previous group is healthy and the comparison shows no unexpected regression. Pause for unexplained changes in latency or timeouts as well as growing pending work or disk use. Check CPU and memory alongside the availability of the monitoring data used to judge progress.
  7. Once all nodes run Cassandra 6.0 execute nodetool cms initialize on one node to complete the metadata transition. Schema changes remain prohibited until it succeeds together with bootstrap and decommission as well as node moves and replacements and assassinate. If initialization reports a mismatch follow that version’s NEWS.txt recovery procedure including the documented stop-and-retry sequence and ignore option.
  8. Run nodetool cms reconfigure to expand the initial one-member CMS to at least three members across the intended failure domains. Schedule storage compatibility changes and optional features separately after the upgrade has been validated.

Optional features

Once the cluster is upgraded you can test optional features one at a time against the workload already validated. Give Accord and compression separate changes from cursor compaction or direct I/O so their effects can be compared. Repair policy and SAI need the same approach alongside guardrails and JVM changes.

ChangeFirst production useCheck
Accord transactionsTest one application workflow with defined transaction boundaries, contention, timeout, retry, and recovery behaviour.Correct committed results, latency under contention, failure recovery, driver support, and operational diagnostics.
Zstd dictionary compressionTest one representative table and dictionary lifecycle.Compression ratio, CPU, flush and compaction rate, read latency, memory use, dictionary distribution, and rollback behaviour.
Cursor compactionEnable only for eligible tables after source SSTable, index, schema, and partitioner checks.Cursor-path selection, allocation rate, compaction throughput, GC, query latency, and fallback behaviour.
Direct I/O modesChange reads or compressed background writes one setting at a time.Page-cache impact, device queueing, direct-write buffer use, compaction rate, and read/write tail latency.
Automated repairStart with conservative assignment and concurrency policy.Repair freshness, history consistency, retries, streamed bytes, compaction pressure, disk headroom, and application impact.
SAI and constraintsTest an explicit query or data-validation requirement.Index build state, selectivity, write cost, query latency, rejected-mutation handling, migration compatibility, and a safe rollback.
JDK 21 and Generational ZGCCompare the existing JVM and collector under identical workload and host conditions.End-to-end latency, throughput, allocation, pauses, concurrent GC CPU, process RSS, direct memory, agents, logs, and restart behaviour.

If latency increases after one feature is enabled you can compare that change with the previous test and configuration. Changing the Cassandra version and storage settings together with drivers and repair policy would leave several possible causes to investigate at once.

When to pause the rollout

Before starting agree on the measurements that will pause the upgrade and who can make that decision. Choose thresholds the team can apply while requests are still running and record the cluster state before making further changes. The following checks give you specific conditions to include in that discussion.

If this happensCheck firstThen
Sustained p99 read or write latency regression beyond the agreed SLOCoordinator and replica latency separately, timeouts, retries, request shape, compaction, disk queueing, GC, and recent deployment event.Pause the rollout and compare the affected node or rack to the preflight baseline before changing more variables.
Timeout or unavailable rate risesConsistency level, replica availability, driver retries, node health, internode errors, and topology state.Stop expanding the change, protect application traffic, and follow the documented recovery or rollback path.
Repair or compaction backlog grows unexpectedlyPending work, table write rate, disk free space, device latency, repair assignment state, and host CPU.Reduce operational pressure, stop the rollout, and determine whether the binary, configuration, or existing capacity caused the backlog.
Schema or metadata disagreementSchema/version state, CMS or topology evidence, node logs, recent DDL, and mixed-version inventory.Halt schema and topology changes, preserve evidence, and use the documented recovery procedure for the Cassandra 6.0 release.
Monitoring or backup coverage is lostMissing metrics/logs, failed agent or JMX connections, backup status, restore evidence, and alert state.Do not continue with an unsupported operating model; restore diagnostic and recovery coverage first.
Disk headroom becomes unsafeFree space by directory, streaming, compaction output, repair work, retention, and write rate.Halt work that expands disk use, recover capacity, and re-evaluate the rollout schedule.

Check how far you can roll back before the upgrade because a new storage format or compatibility mode may prevent a simple return to the previous version. Schema changes and application features can also change what the old release can safely handle. Keep optional features disabled until the upgrade is validated and test rollback or replacement with the drivers and operational tools you depend on. Include monitoring and backups together with repair and topology recovery in that test.

Production acceptance

Keep observing the upgraded cluster through normal workload peaks and an appropriate full repair and compaction cycle before closing the upgrade. Include backup and restore checks with evidence from node restart or replacement tests carried out at supported stages. Review the monitoring used throughout the process so the team can still investigate problems after the rollout ends.

Acceptance areaMinimum proof
AvailabilityThe application meets its agreed availability and latency objectives across all data centers and node groups.
Data protectionRepair freshness is within the required window, backup succeeds, and a representative restore has been verified.
Storage healthDisk headroom, compaction, SSTable counts, tombstones, and page-cache or device behaviour are stable for the workload.
Topology and schemaNodes agree on required metadata, CMS initialization has completed, the CMS has at least three appropriately placed members, topology operations are understood, and application schema changes are controlled.
Client compatibilityEvery production service, job, and administrative tool is operating through a supported and observed driver path.
OperationsMetrics, logs, events, configuration history, alerts, JMX, profiling, backup, and incident runbooks work with Cassandra 6.0.
RecoveryThe team has exercised and documented the rollback, replacement, and failure response boundaries that remain available.

AxonOps keeps Cassandra and host metrics with logs and events alongside configuration and topology history during the upgrade. You can compare an upgraded node with the others while following repair state and changes in the workload. Retaining the before-and-after records lets you check when a regression started and which operations changed at the same time.

Contributors

The series credits contributors on the relevant tickets and pull requests for the changes to transactions and metadata as well as the storage and repair paths. CQL validation and JVM compatibility depend on further work alongside diagnostics and the operational tools used during an upgrade. Test infrastructure and package maintenance are part of getting that work released together with documentation and security review.

I’m grateful to everyone contributing those changes and to the people testing compatibility and investigating difficult failures before and after a release. Running CI and preparing releases takes sustained effort alongside helping users whose production reports reveal cases the tests haven’t yet covered.

Series

Sources

All articles