All articles

Apache Cassandra 6.0 Part 1 - Notes from Using Cassandra Since 2008

Cassandra 6.0 from a long-term user

I have been using Cassandra since 2008 and one of the hardest parts of getting started was finding enough material to learn from. Books and detailed production write-ups were scarce so a problem could send you through mailing lists and Jira discussions before you found something useful to try. Reading the source and testing against a cluster helped fill in the gaps that the available documentation couldn’t cover.

There is far more help available now thanks to the community that has grown around Cassandra over those years. You can find explanations of the architecture alongside practical accounts of running it and use that knowledge when working through your own data model. Understanding partitions and replicas still takes care because the choices you make about consistency and failure domains affect how the cluster behaves. Compaction and repair need the same attention as the data grows and tombstones accumulate.

Cassandra 6.0 gives us plenty to explore across transaction coordination and cluster metadata as well as the storage engine underneath them. Several improvements reduce the work involved in flushing and compaction while others add options for compression and indexing. This series works through those changes with examples so you can see how they fit the Cassandra workloads you already run.

Release status

These posts follow the Apache Cassandra cassandra-6.0 branch and its associated Jira issues and CEPs as reviewed for this update. The line is at 6.0-alpha3 with no beta or release candidate yet while Cassandra 5.0.9 remains the current GA release. You’ll be reading about pre-release behaviour that may change before Cassandra 6.0 is ready for production.

Following the branch gives you time to try the features and understand what an upgrade will involve for your applications. Before planning a production change you’ll still need the final release notes and upgrade documentation for the version you intend to install. Check its compatibility requirements and support status against your cluster as part of that preparation.

What the series covers

AreaWhy it deserves attention
Accord transactionsMulti-partition transactional coordination changes application design options and needs clear contention, timeout, recovery, and driver tests.
Performance and compactionDirect I/O, allocation reductions, flush work, cursor compaction, SSTable writes, and fsync behaviour can affect CPU, GC, page cache, storage queues, and query tail latency.
Repair and diagnosticsA built-in repair scheduler, guardrails, slow-query records, profiling, and richer metrics help operators make repair and failure conditions more explicit.
Zstd dictionary compressionDictionary lifecycle, compression ratio, CPU, memory, SSTable rewrite behaviour, and read cost need table-specific testing.
Transactional Cluster MetadataTCM and CMS replace gossip propagation for critical metadata ordering and make topology and schema state more inspectable.
SAI and CQL constraintsNew index and constraint capabilities extend modelling options, but query selectivity and write-path rejection behaviour still need disciplined validation.
JDK 21The Java 21 profile and Generational ZGC affect heap, direct memory, GC CPU, latency, agents, logs, and restart behaviour.
Upgrade validationA binary upgrade, a storage compatibility advance, and optional feature adoption have different rollback boundaries and should not be combined.

We start with transactions and application behaviour before working through the storage engine and the day-to-day operation of the cluster.

The series

  1. Apache Cassandra 6.0 Part 2 - Accord Transactions
  2. Apache Cassandra 6.0 Part 3 - Performance Optimisations
  3. Apache Cassandra 6.0 Part 4 - Repair, Guardrails, and Observability
  4. Apache Cassandra 6.0 Part 5 - Zstd Dictionary Compression
  5. Apache Cassandra 6.0 Part 6 - Transactional Cluster Metadata and CMS
  6. Apache Cassandra 6.0 Part 7 - Cursor Compaction and SSTable Writes
  7. Apache Cassandra 6.0 Part 8 - Storage-Attached Indexing and Schema Constraints
  8. Apache Cassandra 6.0 Part 9 - JDK 21 and Generational ZGC
  9. Apache Cassandra 6.0 Part 10 - Upgrade and Production Validation

Testing Cassandra 6.0

  • Record the application workload and schema together with the drivers and configuration in use. Include the topology and storage profile as well as the repair schedule and monitoring setup. Keep the recovery procedure with those records so the test includes how you would restore service.
  • Test Cassandra 6.0 before changing storage compatibility or adopting optional features. Keep a supported JDK and the same drivers in place wherever possible so you can identify which change affected the workload.
  • Give each optional feature a separate test with a clear reason for enabling it and a plan for returning to the previous configuration.
  • Run reads and writes while compaction and repair are active and include backups and streaming in the test. Check old SSTables and large or tombstone-heavy partitions as well as node restarts and the tools you use during an incident.

A small synthetic test can help you understand an individual improvement before you try it with several operations running together. You’ll also need that combined workload to see what happens when application traffic competes with maintenance work on the same nodes.

Cassandra and AxonOps

We’ve built AxonOps around operating Cassandra and we also want to make it easier for other people to work with the database. AxonOps Workbench is our open-source desktop client for Cassandra developers and DBAs while CQLAI provides a terminal application with AI assistance. Our practical Cassandra documentation shares guidance for running the database and working through the problems that come with it.

Cassandra has been a large part of our working lives and those projects are some of the ways we give something back. As 6.0 adds more detail about metadata and repair you’ll have more information to examine alongside compaction and query behaviour. AxonOps keeps that information with host metrics and the cluster’s logs and events so you can look back at what happened with the configuration and topology history available too.

The Adaptive Regulation of Cassandra Repair article works through how AxonOps uses high-resolution Cassandra and Linux telemetry to adjust repair velocity and parallelism. It explains how that regulation helps repair finish within the required gc_grace_seconds window while responding to the load on the cluster.

Acknowledging the project

I am grateful to the Apache Cassandra committers and contributors who keep putting their time into the project. A release depends on far more than the features we can point to in a changelog. Someone has to design and review each change and write the tests that will catch a regression later. Others keep CI running and investigate benchmark results or prepare the packages and compatibility guidance that make an upgrade possible. Documentation and proposal discussions need the same care as issue triage and communication with users who are trying to explain a difficult production failure.

The individual posts credit the people named on the associated tickets and pull requests wherever those records identify their work. I also want to thank everyone whose contribution is harder to attach to a single feature because the testing and release work behind Cassandra benefits all of us who run it.

Sources

All articles