All articles

Apache Cassandra 6.0 Part 9 - JDK 21 and Generational ZGC

JDK 21 and Generational ZGC

Changing the JDK on a Cassandra node can affect the latency you see even when its data and query workload stay the same. The collector and JIT are part of that change alongside memory layout and the logging and native-library behaviour used by the process. You’ll also need to check module access and security-manager handling with the agents and JMX tools used to operate the node.

Cassandra 6.0 supports Java 11 and 17 alongside Java 21 while keeping Java 11 as its default build JDK. The build file marks builds using non-default JDKs as experimental and intended for development and testing. That makes it worth checking the exact build and runtime combination as well as the collector and heap settings you intend to use.

This post follows pre-release Cassandra 6.0 as reviewed against 6.0-alpha3 and the details may change before GA. Check the release notes and upgrade documentation for the version you’re testing before planning a JVM change.

AreaCassandra 5.0 baselineCassandra 6.0 work to validate
Runtime versionsJava 11 and 17 are the established supported options for most Cassandra 5.0 deployments.The build accepts Java 21 and supplies a Java-21-specific server options file.
Default collector profileExisting deployments commonly run an established collector and a locally tuned heap.The JDK 21 file enables ZGC with its generational mode and provides an alternative G1 configuration as comments.
Memory-layout settingsHeap, direct memory, compressed references, and host-page configuration are usually tuned locally.Cassandra carries a temporary -XX:-UseCompressedOops workaround for the JDK 21 ZGC and Jamm interaction.
GC loggingGC logs can introduce file-I/O work on the JVM path.Async log mode is enabled in the JDK 17 and 21 option files to reduce log-file I/O stalls, with an explicit loss mode if its buffer fills.

What Cassandra configures for JDK 21

Cassandra selects jvm21-server.options for Java 21 and newer so start by checking the active settings in that file.

-XX:+UseZGC
-XX:+ZGenerational
-XX:-UseCompressedOops
-Xlog:async

The first two flags enable Generational ZGC while the third works around Jamm’s incorrect default assumption about compressed ordinary object pointers in this configuration. Keep that workaround in place because Jamm uses the memory-layout information when calculating object sizes. The fourth flag changes logging by buffering records in memory and writing them separately so GC log-file I/O is less likely to block a JVM thread.

You’ll also find commented G1 settings and options for explicit large pages and SoftMaxHeapSize in the file. Further options control whether ZGC uncommits unused memory and how long it waits before doing so. These are alternatives to evaluate against the host and workload rather than settings that are already active.

SettingBehaviourValidation required
-XX:+UseZGCEnables ZGC, a concurrent low-latency collector.Compare end-to-end latency and application throughput with the current collector on the same workload.
-XX:+ZGenerationalSeparates young and old collection work so short-lived allocation can usually be reclaimed without scanning the older live set.Check allocation rate, concurrent GC CPU, generation behaviour, heap headroom, and latency under compaction and repair.
-XX:-UseCompressedOopsAvoids the Jamm compressed-reference measurement issue documented in Cassandra’s JDK 21 options.Calculate heap and process memory requirements again; pointer size affects the retained-heap footprint.
-Xlog:asyncBuffers JVM log records before they are drained to the configured output.Confirm the log pipeline, free space, rotation, and that the asynchronous buffer does not lose records under the expected logging volume.
-XX:+AlwaysPreTouch (commented)Faults heap pages at startup instead of allowing later page faults during normal traffic.Measure restart time, required memory availability, and whether the latency profile improves after startup.

Cassandra continues to default to G1 on JDK 11 and 17 while the JDK 21 profile selects Generational ZGC. On JDK 21 keep -XX:+UseZGC together with -XX:+ZGenerational because Cassandra doesn’t recommend non-generational ZGC. The flag changes on later Java releases with ZGenerational deprecated in JDK 23 and non-generational ZGC removed in JDK 24 so check the supported runtime you’re actually using.

How Generational ZGC works

Consider a Cassandra thread loading a reference while the garbage collector moves the referenced object elsewhere in the heap. Generational ZGC retains ZGC’s concurrent compacting design and uses checks called barriers on reference loads and stores to keep that access correct. It can do the remaining reference work while Cassandra runs instead of stopping the application for a heap-wide pointer update.

A compacting collector moves live objects out of sparsely occupied regions so their space can be reused without leaving the heap fragmented. The old address can remain in a field while relocation is in progress so the next thread reading it needs a way to find the object. ZGC records GC state in references and checks that state when Java code loads an object reference.

Coloured references and barriers

A ZGC reference combines the object address with a small amount of GC state that the load barrier checks against the current collection phase. If its mark or remap state is already current the check takes a small register operation before the thread can continue.

If the states differ the load barrier uses a slower path to bring the reference up to date. During concurrent marking it can mark the object and during relocation it uses the forwarding table to find the new address. It returns a current reference and can repair the field with compare-and-swap so a later load can take the fast path.

Now consider a retained cache or schema object being updated to point at a newly allocated young object. A young collection needs to find that reference even though it isn’t tracing the whole old generation. Generational ZGC’s store barrier records the field in a remembered set so the next young collection can use it as a root alongside thread-stack references.

The store barrier also preserves the reference being overwritten when snapshot-at-the-beginning marking is in progress. Recording it before the write completes lets the marker retain a path to an object that was live when marking began even if the application replaces its last reference.

A young collection

A young collection starts with a short pause to update shared GC state including the current mark state and remembered-set snapshot. The pause doesn’t walk the Java heap and concurrent marking then follows the live young objects from thread stacks and remembered references while Cassandra continues running.

After marking identifies the live objects ZGC selects sparse young regions for relocation and installs forwarding tables. It copies the surviving objects into new regions and returns the old regions to the allocator while load barriers forward and repair references as they are used. Adaptive tenuring determines which survivors move into old space where less frequent collections use the same broad marking and relocation approach.

Memtable data is included in the Java heap when Cassandra uses its default heap_buffers allocation mode. Memtables participate as live heap data without forming a separate ZGC generation while offheap_* modes move part of their storage outside the heap. Direct buffers and SSTables remain outside the Java heap along with the operating-system page cache.

Cassandra performance impact

To judge the effect on Cassandra first check how much of the delay comes from allocation and heap collection. A read waiting on an SSTable fetch or a replica response may see little improvement if storage or network queues dominate its latency. Generational ZGC changes the collection work around those requests so its benefit depends on whether that work is limiting the node.

Request decoding and CQL execution create short-lived objects as Cassandra builds results and serializes responses. Internode messaging and maintenance add more temporary allocation while caches and prepared statements retain objects for longer. Schema metadata and executors stay live too alongside metrics and the active node’s data structures. Frequent young collections can focus on temporary objects without repeatedly tracing that whole retained old set.

On a node with high allocation pressure you may see fewer GC-related spikes in the slowest requests. Short stop-the-world phases reduce one source of delay and young collections avoid repeatedly tracing the entire old generation. You’ll still need to measure allocation stalls because requests can wait if memory isn’t reclaimed quickly enough and load barriers can take their slower path during marking or relocation.

Oracle’s Java team published a Cassandra example in Inside Java’s Generational ZGC explainer that helps illustrate this behaviour. Single-generation ZGC began suffering allocation stalls above 75 concurrent clients while Generational ZGC maintained its pause profile through 275 clients in that test. Those counts belong to the benchmark setup and should be checked against your own allocation rate and heap size before drawing a capacity conclusion. CPU resources and the data model affect the comparison too alongside driver behaviour and storage latency.

Cassandra measureWhat Generational ZGC can changeWhat still requires measurement
Coordinator and replica p95, p99, and p99.9 latencyShort GC pauses can remove a node-wide pause contribution. Lower allocation-stall risk can prevent isolated requests from waiting for space to be reclaimed.Disk reads, page-cache misses, queueing, replica load, network delay, speculative execution, and timeouts can dominate the same percentiles. Capture client and server-side latency together.
Read and write throughputYoung collections can spend less time tracing retained old objects, leaving more CPU time for useful work when allocation pressure was previously high.Java reference load and store barriers have a CPU cost. The store barrier is triggered by Java reference writes, not by a CQL write as such. Measure sustained operations per second with the same data, consistency level, and concurrency.
Node CPULess collection work per short-lived allocation can reduce GC pressure. Avoiding allocation stalls can also avoid a collapse in useful request work at high concurrency.Concurrent GC workers run beside Cassandra executors, compaction, repair, compression, encryption, and network processing. A lower pause graph can coincide with higher total CPU use. Track process CPU and GC-worker CPU, not just pause duration.
Heap capacity and process RSSGenerational ZGC adapts generation sizes and tenuring thresholds to the allocation rate and live set. A well-sized heap gives concurrent collection room to run while requests continue.Cassandra’s JDK 21 options disable compressed ordinary object pointers for the Jamm workaround. Equal -Xmx values can therefore produce a different retained-heap and resident-memory footprint. Recalculate host and container headroom.
Allocation stallsFrequent young collection reduces the chance that short-lived allocations force work against the full retained heap.A node can still stall when allocation exceeds reclamation capacity. Treat any allocation-stall event as an application-latency investigation, even when stop-the-world pauses remain small.
Compaction, repair, and streamingA collector with a steadier latency profile can reduce Java-heap interference while maintenance runs alongside traffic.These operations remain constrained by disk bandwidth, page-cache residency, network throughput, compaction strategy, and the CPU they share with the collector. GC cannot compensate for an overloaded storage path.

If p99 follows page-cache misses and device queueing on a read-heavy node the collector may have little visible effect. A coordinator with high concurrency and allocation may benefit much more if its worst requests coincide with pauses or allocation stalls. Compare the request timeline with both sets of measurements so you can tell which delay changed.

Run normal client traffic while compaction and flushing are active and repeat the test during repair or streaming. Include bootstrap and replacement where they’re part of the workload you need to support. Compare latency and throughput with CPU and resident memory so an improvement in GC behaviour doesn’t consume memory the page cache needs.

Heap, direct memory, and the host

Giving JDK 17 and JDK 21 the same -Xmx doesn’t tell you whether they leave the same memory available to the host. The process also uses direct buffers and native allocations together with thread stacks and JIT code while libraries occupy further space. The operating system needs room for page cache and filesystem metadata as well as the kernel so compare the complete resident footprint.

The options for SoftMaxHeapSize and uncommit are commented because their effects depend on the host and the workload. Uncommitting unused memory can reduce process footprint but recommitting and faulting those pages later may add latency. Setting -Xms equal to -Xmx with AlwaysPreTouch can avoid later heap page faults by reserving and touching the memory at startup. Test both the startup cost and steady-state footprint before choosing the approach that fits the node.

The options file leaves UseLargePages disabled so explicit huge pages need a deliberate host configuration before testing. They may help throughput and latency on a suitable host while transparent huge pages can introduce latency spikes in a sensitive workload. Keep the page settings with the results so two JVM runs don’t silently use different memory behaviour.

Memory areaEvidence to captureCommon mistake
Java heapUsed heap, committed heap, allocation rate, live-set estimate, GC cycles, pauses, and concurrent-GC CPU.Concluding that low used heap means the node has spare memory.
Direct and off-heap memoryDirect-buffer use, native memory tracking where available, allocator metrics, and process RSS.Keeping the old direct-memory limit after changing heap or collector settings.
Page cacheMajor faults, cache residency, device reads, and query latency through compaction.Giving nearly all RAM to the JVM and starving the cache used for SSTable reads.
Container or cgroup limitLimit, current use, OOM events, memory pressure, and host-level eviction evidence.Treating -Xmx as the complete process memory budget.
Startup memory policyStartup duration, resident memory after start, page faults under first traffic, and recovery time.Enabling pre-touch without reserving enough capacity for rolling restart overlap.

Async GC logging

With CASSANDRA-21372 Cassandra enables -Xlog:async on JDK versions that support it. GC-related log records go into a circular memory buffer that is drained separately so a slow log-file write is less likely to block a JVM thread.

CASSANDRA-20980 adds separate GCInspector thresholds for concurrent GC events which need to be interpreted alongside pauses. Record concurrent-GC CPU and allocation stalls too so a low pause time doesn’t hide the rest of the collector’s work.

If log records arrive faster than the asynchronous buffer can drain Cassandra’s option file warns that records are silently dropped. Check whether the log destination can keep up before increasing the buffer and verify disk capacity and rotation with the log-shipping and retention setup. Buffering can reduce logging stalls but it doesn’t reduce allocation rate or make the collector itself faster.

Compare asynchronous logging with the same log settings and traffic while recording log volume and any signs of dropped records. Check the destination’s storage latency against GC pauses and allocation stalls alongside the latency seen by the application. Keep JVM metrics and profiler results with the logs so missing diagnostic records don’t leave the comparison incomplete.

Cassandra 5.0 and 6.0 validation

Keep the Cassandra version and schema unchanged during each collector comparison along with data volume and replication settings. Use the same host class and compaction configuration with the same client driver workload before changing another setting. Combining the JDK move with a different heap or storage device would make it harder to explain which change affected the result.

TestBaselineJDK 21 comparisonMeasurements
Steady client trafficCurrent supported JDK and production collector configuration.JDK 21 with Cassandra’s supplied options, then a controlled heap change only if needed.Throughput, p50/p95/p99 coordinator and replica latency, retries, timeouts, allocation rate, GC pauses, concurrent-GC CPU, and RSS.
Compaction and flushSame active read and write workload while maintenance runs.Repeat with JDK 21, preserving compaction and flush concurrency.Read and write tail latency, page faults, device queueing, compaction throughput, heap and direct memory, and GC activity.
Repair and streamingExisting repair workload with realistic data size and network use.Repeat on JDK 21 during repair, bootstrap, replacement, or streaming.Session completion, latency, CPU, off-heap memory, GC logs, network throughput, and disk headroom.
Restart and recoveryCurrent JVM with the existing startup policy.JDK 21 with and without any deliberate pre-touch or uncommit choice.Startup duration, resident memory, first-request latency, page faults, time to become healthy, and impact on other nodes.
Observability and agentsExisting JMX, metrics, tracing, backup, profiler, and log collection path.Repeat every operational workflow against JDK 21.Successful attachment, metric continuity, agent errors, log format and retention, backup/restore behaviour, and alert compatibility.

Record the JDK vendor and patch level because the major version alone doesn’t identify the runtime used in a test. Vendor and patch differences can affect TLS and native libraries alongside GC fixes and the tools attaching to the process. Keep JVM flags and cassandra-env settings with cgroup limits and kernel details as well as page configuration and agent versions.

In AxonOps you can compare JVM and GC measurements with Cassandra requests and the compaction or repair running at the same time. Disk I/O and configuration history help explain the result alongside logs and events from each run. That lets you check whether the cluster handled its normal workload better while retaining enough information to investigate later problems.

Contributors

Achilles Benetopoulos proposed JDK 21 support with Josh McKenzie assigned to the work while Dmitry Konstantinov proposed and implemented asynchronous GC logging for the JDK 17 and 21 options. Supporting those paths also takes testing across multiple JDKs and investigation of compatibility failures in dependencies and agents. Review of JVM flags and packaging needs the same care as documenting what changes for operators.

I’m grateful to everyone doing that work because every Cassandra node depends on it when starting and recovering as well as handling requests. Testing the runtime and its operational tools gives us a better basis for both normal maintenance and incident investigation.

Series

Sources

All articles