All articles

Apache Cassandra 6.0 Part 7 - Cursor Compaction and SSTable Writes

Cursor Compaction and SSTable Writes

Part 3 followed compaction reads and memtable flushing while this post looks at how Cassandra performs the compaction merge itself. It has to reconcile partitions and cells from several SSTables and remove data only when safe before writing the replacement files.

This post follows pre-release Cassandra 6.0 as reviewed against 6.0-alpha3 so these details may change before GA. Check the release notes and upgrade documentation for the version you’re testing before changing compaction settings.

The cursor implementation in CASSANDRA-20918 reduces the temporary Java objects created while reading and merging SSTables. Related work adds direct I/O for compressed background writes through CASSANDRA-21134 and connects cursor reads to the compaction access mode through CASSANDRA-21147. The expanded compaction history in CASSANDRA-20081 gives you more information to retain when comparing those paths.

If a node spends much of its time compacting then reducing work in that merge can leave more resources for requests. Cursor compaction remains experimental and is disabled in cassandra.yaml while cassandra_latest.yaml enables it. Test the cursor implementation separately from direct background writes and trickle fsync because each changes a different part of storage maintenance.

ChangePotential benefitCondition to verify before using it
Cursor compactionMuch lower compaction allocation rate and less CPU spent constructing short-lived row and cell objects.The table, source SSTable format, partitioner, schema, and compaction operation are supported by the cursor path.
Direct compaction readsLarge sequential reads do not displace hot query pages from the operating-system page cache.Query latency, device queueing, and compaction throughput improve together on the target storage.
Direct background writesCompressed SSTable output from maintenance work avoids filling the page cache.The lower cache pressure is worth the direct-I/O buffering and device behaviour on that workload.
Trickle fsyncBuffered output is synchronized in bounded intervals instead of accumulating into a larger writeback burst.The chosen interval protects read latency without reducing maintenance throughput too far.

Cassandra 5.0 and 6.0

Cassandra 5.0 introduced TrieMemtable and the trie-indexed bti format alongside Unified Compaction Strategy. Those changes affected how data is held in memory and stored on disk as well as which SSTables are selected for maintenance. UCS also gave users more control over the balance between read and write amplification and space usage.

During an ordinary compaction Cassandra decodes the selected SSTables into its usual in-memory objects before merging their contents. Partition keys and clustering values need representations alongside liveness and deletion information for the rows and cells being processed. Many of those objects survive only until the entry is serialized into the output SSTable but allocating them still uses CPU and creates young-generation collection work.

Cursor compaction uses reusable descriptors to read and merge SSTable contents before writing the result without building the usual per-entry object graph. The compaction strategy still chooses the SSTables and the merge still follows the same reconciliation and tombstone-safety rules. The change is in how Cassandra executes the selected work which lets you compare the two implementations under the same strategy.

AreaCassandra 5.0 baselineCassandra 6.0 work
In-memory data before flushTrieMemtable reduces memory and garbage-collection costs while mutations are resident.The flush path receives specialized serialization and metadata work, covered in Part 3.
Compaction selectionUnified Compaction Strategy and the established strategies decide which SSTables need maintenance.The same strategy can use a cursor merge path when the table and operation meet its eligibility checks.
Merge executionThe iterator path materializes normal Cassandra row and cell structures while it reconciles input SSTables.Reusable cursor descriptors parse and merge the SSTable stream with far less short-lived allocation.
Compaction I/OBuffered reads and writes use the page cache unless another disk mode is selected.Direct read mode is available to cursor compaction; compressed background writes can also use direct I/O.

Before interpreting a benchmark you’ll need to check whether the tested compactions actually used the cursor path. Eligibility depends on the table and operation so enabling the setting alone doesn’t show how much of your workload will benefit.

Cursor compaction mechanics

The input SSTables are sorted by partition key and clustering order which lets Cassandra merge corresponding entries across the files. It selects the latest live version of each item while retaining tombstones that can’t yet be safely removed. The output stays sorted as it is written and the cursor implementation aims to do that work with less temporary allocation.

You can follow the implementation through SSTableCursorReader into CursorCompactor and then SSTableCursorWriter in the Cassandra source. The reader advances through partition starts and static rows before processing rows and their cell headers and values. It also recognizes range tombstone markers and partition ends as part of that state machine. Reusable PartitionDescriptor and UnfilteredDescriptor objects hold the current values with liveness and deletion-time and cell structures so advancing the reader can reuse their storage.

The compactor keeps one cursor per input SSTable and orders them by their current partition and clustering positions. It finds entries for the same logical value and applies Cassandra’s timestamp and deletion rules before asking the writer to emit the surviving data. The writer serializes SSTable components and updates their indexes and metadata while using bounded scratch buffers where the format needs a complete row before it can write headers or complex-column markers.

Iterator compaction
The bar shows allocation accumulated over the compaction rather than memory held at one time because garbage collection can reclaim discarded objects while compaction continues.
Cursor compaction
The boxes show a reusable holder at successive points during a file scan with one cursor per input SSTable. Object reuse reduces allocation and GC work while the input files still need to be read and the surviving data written into replacement SSTables.
StepConventional iterator pathCursor path
Read a partitionDecode the partition into standard objects used by the iterator pipeline.Advance an SSTableCursorReader and load the current partition descriptor into reusable storage.
Read rows and cellsConstruct row, cell, liveness, and deletion representations as the stream is traversed.Advance through reader states such as row start, cell header, and cell value while reusing descriptors.
Reconcile input SSTablesMerge iterator entries and apply deletion and liveness rules.Sort the active cursors, merge matching descriptors, and apply the same rules without retaining a per-entry object graph.
Write outputSerialize merged objects into a new SSTable.SSTableCursorWriter writes the surviving descriptors and updates the BIG-format index and metadata.
Move to the next partitionDiscard temporary objects and repeat.Reuse the same descriptor and scratch objects for the next partition.

The early benchmark suite in CASSANDRA-20918 reported around 20 MB of heap allocation for cursor compaction against more than 5 GB for the regular implementation in the tested cases. Several test mixes also showed compaction running three to five times faster with the cursor path. Those development results show why reducing temporary objects is worth testing but your own throughput and GC savings depend on the workload.

The cursor writer has to preserve row indexes and partition metadata alongside checksums and the serialized structure of simple and complex columns. Its lower allocation rate is useful only if the same values win reconciliation and the same tombstones remain afterwards. That is why replacing the iterator path needs tests of output correctness as well as throughput.

Eligibility and fallback

Check which configuration profile you’re using because cursor_compaction_enabled is false in cassandra.yaml and true in cassandra_latest.yaml. Both describe an experimental garbage-free compaction path that needs production-like testing before you rely on it.

# cassandra.yaml
cursor_compaction_enabled: true

Even with the flag enabled Cassandra checks each compaction and selects the iterator path for unsupported tables or operations. That fallback lets the maintenance work proceed correctly while leaving the allocation optimization unused for that particular job.

Alpha3 adds non-frozen collection and UDT support through CASSANDRA-21463 while counter tables still use the iterator path. An SSTable header containing a dropped multi-cell column also triggers fallback so schema history affects eligibility. The alpha releases found output differences involving index-offset overflow and same-timestamp tie breaks as well as dropped-column filtering and materialized-view row resurrection. CASSANDRA-21462 and CASSANDRA-21255 cover fixes alongside CASSANDRA-21152 and help explain why the implementation remains experimental.

SituationCursor-compaction behaviourOperational implication
Plain supported table and current BIG SSTablesCursor compaction can run when enabled.Establish the allocation, CPU, compaction-throughput, and query-latency baseline before enabling it broadly.
Table with countersCassandra uses the iterator path.A cluster can contain both paths. Do not assume the setting changes every table’s maintenance profile.
Non-frozen collections and UDTsCursor compaction can run when the other eligibility checks pass.Alpha3 removed the earlier collection fallback, but this still needs schema-specific tests.
Dropped multi-cell column retained in an SSTable headerCassandra uses the iterator path.Test tables with the schema history and SSTable generations that exist in production, not only a newly created table.

Use each table’s existing SSTables when testing eligibility because the current CQL schema may hide older columns recorded in their headers. Include a mixture of active tables and old files while normal compaction and repair-related maintenance run alongside requests. A freshly compacted test table won’t exercise all the combinations present on a long-running cluster.

Direct I/O for SSTable reads and writes

The cursor reader uses the configurable SSTable access mode connected by CASSANDRA-21147. Setting compaction_read_disk_access_mode to direct lets a supported cursor compaction bypass the operating-system page cache for its compressed SSTable input reads just as the iterator path can.

Consider compaction scanning a large input set while queries repeatedly read data that normally stays in the page cache. Bypassing the cache for compaction input can prevent that scan from evicting the pages those queries need. Client reads keep their existing access path and both workloads still compete for the storage device.

For background output Cassandra 6.0 provides background_write_disk_access_mode to control writes to compressed SSTables. It covers compaction and streaming together with cleanup and repair as well as SSTable upgrades.

# cassandra.yaml
# Applies only to compressed SSTable writes from background operations.
background_write_disk_access_mode: direct

# Per concurrent background writer. This uses off-heap memory.
direct_write_buffer_size: 1MiB

The implementation described in CASSANDRA-21134 keeps uncompressed writes buffered and leaves memtable flushes buffered too. Recently flushed data can then remain in the page cache for subsequent queries while direct background writes use staging buffers aligned to filesystem requirements. Each concurrent background writer gets its own buffer so increasing compaction or streaming concurrency also increases that off-heap memory use.

Direct background writes can help keep query data in the page cache but give up the kernel buffering that normally absorbs and schedules those writes. Run the test with concurrent reads and writes while compaction processes enough data to exceed memory. You’ll need to see whether query latency improves and maintenance still keeps up under that combined load.

Trickle fsync

For buffered writes Cassandra 6.0 enables trickle_fsync with a default trickle_fsync_interval of 10240KiB. Cassandra calls fsync() periodically during sequential writes so less dirty data can accumulate before a forced writeback.

The interval counts uncompressed bytes so a compressed SSTable may trigger fsync after writing fewer physical bytes than that value suggests. A smaller interval can limit writeback bursts and help read tail latency while increasing fsync frequency and potentially reducing write throughput. Raising it reduces that overhead until the kernel’s ordinary dirty-page writeback starts determining the timing again.

Storage conditionWhat to compareSignals that decide the setting
SSD or NVMe with concurrent readsDefault interval against the intended change under the same compaction load.p95/p99 read latency, write latency, device queue depth, fsync duration, compaction progress, and host CPU.
Network-attached or cloud block storageBuffered and direct background writes, with the same compaction concurrency.Tail latency during storage stalls, IOPS and throughput limits, write throttling, and compaction completion time.
Spinning disksDefault trickle fsync against a carefully larger interval.Read latency variability, seek pressure, throughput, and whether maintenance falls behind.
Write-heavy tables that are rarely read after flushBuffered flush behaviour should remain the reference case.Page-cache usefulness after flush, read latency when a hot set exists, and disk behaviour during flush plus compaction.

Compaction throughput and history

Faster compaction can shorten the time SSTables overlap and reduce read amplification while also changing when later jobs are selected. In the CASSANDRA-20918 review Dmitry Konstantinov reported lower compaction-thread allocation and roughly twice the throughput for a VInt-heavy test. He also noted that faster completion can leave less time for SSTables to accumulate into one batch under intensive writes. Check queueing and the resulting SSTable layout alongside throughput so you can see how the strategy responds.

CASSANDRA-20081 adds compaction type and strategy name together with the level to nodetool compactionhistory. You can retain those details with each event to identify the strategy that produced the result when comparing runs.

UCS can split a compaction into output-shard tasks through parallelize_output_shards so those tasks can run in parallel. Major compactions can use that parallelism through nodetool compact --jobs with the default limited to half the available compaction threads. The remaining threads let ordinary background compaction continue while the major compaction runs.

TestBaselineCassandra 6.0 comparisonMeasurements
Cursor implementationcursor_compaction_enabled: false with the production table definition and SSTable generations.Enable it only for tables that are eligible, then confirm which path actually runs from logs and profiling.Compaction throughput, compaction-thread allocation rate, CPU, GC activity, pending compactions, SSTables per read, and read p95/p99.
Compaction readsBuffered read access for compaction.compaction_read_disk_access_mode: direct, tested with the cursor path where eligible.Page-cache behaviour, major faults, device queue depth, compaction rate, and read latency during and after compaction.
Compressed background writesbackground_write_disk_access_mode: standard.Direct background writes with a controlled direct_write_buffer_size.Off-heap memory per active writer, device latency, write throughput, compaction completion time, and client read latency.
Trickle fsyncThe supplied interval with the same storage and write concurrency.One controlled interval change at a time.Fsync latency, tail read and write latency, dirty writeback behaviour, device utilization, and maintenance backlog.
Mixed production trafficReads, writes, repair-related maintenance, compaction, and normal operational concurrency.Repeat after each independent setting change.Coordinator and replica latency, table-level SSTables per read, tombstones scanned, heap and off-heap allocation, logs, and full configuration.

Keep the compaction strategy and its options with the schema and SSTable format for each run. Record compression and storage compatibility settings alongside the device and filesystem with the kernel and JDK versions. Include concurrent compactor count and the amount and age distribution of data so another run can reproduce the workload.

In AxonOps you can compare Cassandra compaction and table metrics with host I/O and the latency seen by clients. JVM allocation and GC activity help explain changes in CPU while logs and configuration history show which settings selected the path. Reading those measurements together helps you check whether a delay coincided with page-cache eviction or storage queues as well as changes in SSTable count and compaction backlog.

Contributors

Nitsan Wakart authored cursor compaction with Branimir Lambov and Dmitry Konstantinov credited for review in the merged Apache Cassandra pull request. The ticket also records guidance from Benedict Elliott Smith and benchmark work from David Capwell alongside Josh McKenzie’s help with the contribution process and merge. Testing and profiling continued beyond the microbenchmarks to establish how the implementation behaved under more varied conditions.

Sam Lightfoot reported and implemented the direct background-write work and the cursor direct-read follow-up. Brad Schoening requested the compaction-history additions and Arvind Kandpal implemented them with Maxwell Guo and Jyothsna Konisa credited for review.

I’m grateful to those contributors and to everyone maintaining the compaction tests and investigating failures in CI. Follow-up review and documentation take time alongside release preparation and support for people encountering problems in their own clusters.

Series

Sources

All articles