Repair, Guardrails, and Observability
A replica can return from an outage with writes missing because hints expired or the data changed while it was unavailable. Repair checks for those differences and brings replicas back into agreement before missed deletes have a chance to reappear. The time available depends on tombstones remaining on the replicas that need to take part in that repair.
Cassandra users have long used scripts and external services to decide which ranges to repair and how much work to run in parallel. Running the command is only part of that job because someone also needs to track completion and retry failed work. Disk headroom and topology changes affect the schedule too so the controller needs a way to respond when conditions change halfway through a run.
Cassandra 6.0 adds a built-in repair scheduler alongside guardrails and diagnostic information for investigating problems on a running cluster. We’ll work through how the scheduler assigns work and records progress before looking at the decisions involved in controlling repair under application load.
This post follows pre-release Cassandra 6.0 as reviewed against 6.0-alpha3 so these details may change before GA. Check the release notes and upgrade documentation for the exact version you’re testing before adopting its settings.
| Area | Cassandra 5.0 baseline | Cassandra 6.0 work to validate |
|---|---|---|
| Repair control | Operators commonly coordinate repair with external schedules and retain their own repair history. | CEP-37 introduces a scheduler and replicated repair-history state inside Cassandra for full, incremental, and preview-repaired categories. |
| Disk protection | A disk guardrail can protect a replica or token range as capacity becomes unsafe. | A keyspace-wide option can stop writes across all replicas of a keyspace when one participating node crosses the failure threshold. |
| Client compatibility | Driver inventory and enforcement are usually external release-process work. | A server-side driver-version guardrail can warn or reject a declared driver type and version below a configured minimum. |
| Diagnosis | Logs, JMX metrics, tracing, and external profilers are the usual evidence sources. | Slow-query records, integrated async profiling, administrative history, and richer table metrics give operators more Cassandra-native evidence. |
Automated repair architecture
The Apache Cassandra Unified Repair Solution is the work described in CEP-37 to schedule repair and retain its history inside Cassandra. The underlying repair still compares replica data and streams differences when required while the scheduler decides when and where to run that work.
The scheduler stores replicated repair history in system_distributed so nodes can see what has been repaired and assign follow-up work across the ring. A scheduler thread pool coordinates assignments that can be split into smaller token ranges to keep individual jobs bounded. It supports full and incremental repair together with preview-repaired runs while Paxos repair remains a separate proposal outside the 6.0-alpha3 scheduler.
Follow an assignment through the table below to see how the scheduler decides what can run and records what happened.
| Step | Scheduler activity | Why the detail is operationally important |
|---|---|---|
| 1. Read repair history | Nodes use replicated state to establish which table and token ranges need attention. | Repair selection has an auditable history rather than existing only in a cron job or an external service database. |
| 2. Create bounded assignments | The scheduler divides the work by token range, data size, or partition count according to its configuration. | Assignment size controls the burst of disk, CPU, network, and compaction work introduced by repair. |
| 3. Select eligible work | It considers repair type, node and replica constraints, schedules, retry state, major-version safety, and configured data-center scope. | The system needs to avoid turning an already overloaded node or an incompatible rolling upgrade into more repair work. |
| 4. Run and record repair | Cassandra carries out the selected repair and updates replicated history as it progresses. | Completion must be distinguishable from a job that was started, retried, interrupted, or failed after streaming began. |
| 5. Revisit the ring | The scheduler returns to ranges and tables according to its interval and repair policy. | The useful measure is repair freshness across the full data set, not an isolated successful session. |
You’ll find global settings and per-repair-type overrides in the auto_repair section of the Cassandra 6.0 configuration. Use bytes_per_assignment and partitions_per_assignment with the maximum tables per assignment to control how much work each job takes on. Other settings govern the minimum interval before repairing the same node again together with retry limits and backoff. You can also limit work to primary token ranges or exclude data centres while grouping tables by keyspace and choosing how many nodes may repair concurrently.
The auto_repair section ships commented out so the scheduler won’t run until you configure it. The supplied example enables only full repair and gives you a configuration to test against your own repair requirements.
Consider a cluster whose largest tables take longer to repair than the rest while application writes continue throughout the day. You need to choose assignments and parallelism that let the whole repair cycle finish within its safety window without overloading the nodes. Individual commands can succeed while the overall cycle still takes too long so completion time needs to be tracked across the full set of work. Adaptive Regulation of Cassandra Repair explains how AxonOps adjusts repair velocity and parallelism using live cluster conditions while working towards that deadline.
Repair state and failure handling
Suppose one scheduler removes a repair-history entry while another is trying to update the same record. The resulting history needs to describe the work consistently or a later assignment could be based on an incorrect view of what finished. CASSANDRA-20996 proposes lightweight transactions for all auto-repair history mutations but remains open in 6.0-alpha3 with LWT used for only a subset of those operations.
Using LWT adds coordination to those history updates so the implementation needs to balance that cost against the consistency of scheduler decisions. Test what the history records when a node restarts during a retry or cleanup overlaps another scheduler update. A repair completing successfully in isolation won’t exercise those interactions between nodes updating shared state.
Repair may stream data and perform anti-compaction while temporary files increase disk use on the participating nodes. Cassandra 6.0 can reject repair at a configured disk-headroom threshold and hold back new work when pending compaction exceeds a configured threshold. Size those limits around the largest permitted assignment and measured free space so they take effect before the node runs out of room.
| Failure or pressure condition | Required evidence | Expected operating response |
|---|---|---|
| A node restarts while repair history is updated | History record state before and after the restart, scheduler logs, and repaired-range freshness. | Confirm the scheduler does not duplicate or lose the assignment and that its retry state is explicit. |
| Repair is delayed by compaction pressure | Pending compactions, repair-rejection event, table write rate, and disk queueing. | Reduce repair parallelism or assignment size, address the backlog, then resume with a documented schedule. |
| Disk headroom falls below the repair threshold | Free space by data directory, repair activity, compaction output, and retention trend. | Stop adding repair pressure, recover capacity, and establish why the headroom model was wrong. |
| Rolling major-version upgrade is in progress | Node version inventory, repair configuration, and scheduler events. | Follow the mixed-version guardrail rather than allowing automated work to mask an upgrade issue. |
| Network or replica failure interrupts a session | Session state, retry count, streamed bytes, timeout and error logs, and post-retry repair history. | Verify the retry is bounded and the range remains visible until it has actually completed. |
The built-in scheduler can take care of assignments and retries while you follow the progress and impact of each repair cycle. You’ll still need a plan for what happens when the cluster can’t finish the required work in time.
Repair control under live load
CEP-37 gives Cassandra a native scheduler with replicated history and token-range assignments together with retries and protective thresholds. I’m pleased to see that work in the project because it gives users a built-in option for scheduling repair and understanding what completed.
AxonOps Adaptive Repair uses high-resolution Cassandra and Linux telemetry to adjust repair velocity and parallelism against the time remaining before each table’s gc_grace_seconds deadline. If client traffic increases while a repair is running it can respond to the resulting replica latency and I/O wait. Compaction pressure is part of that feedback too so the repair plan can change as the load on the nodes changes.
| Control question | CEP-37 scheduler | AxonOps Adaptive Repair |
|---|---|---|
| How is work organised? | Cassandra maintains repair history and schedules bounded assignments with repair-type, concurrency, retry, and policy controls. | AxonOps plans token-range segments against the table volume and repair deadline, then continuously regulates their launch rate. |
| What is the operating input? | Configured assignment, interval, concurrency, safety, and retry settings within Cassandra’s scheduler. | Current repair progress plus 5-second Cassandra and Linux telemetry, including pending ReadStage and MutationStage work, coordinator and replica latency drift, and I/O wait. |
| What happens as the cluster becomes busy? | The configured scheduler and protective thresholds govern whether work remains eligible. | The feedback controller can reduce or increase repair velocity and parallelism as the live performance state changes, while preserving the completion target. |
| What is the target? | Reliable automatic repair orchestration within Cassandra. | Complete repair inside the required safety window without treating a live cluster as if it were idle. |
When you’re evaluating repair control it’s worth testing what happens after the application load changes during a run. The Adaptive Regulation of Cassandra Repair article follows the AxonOps feedback model through the telemetry it uses to make those adjustments. It also explains how regulation accounts for uneven table sizes and changing host I/O conditions while repair is in progress.
Disk and driver guardrails
A guardrail lets Cassandra warn about a configured risk or reject an operation before it pushes the database further into that condition. You’ll still need capacity planning and schema review alongside client-release checks and monitoring that gives you time to act before rejection begins.
If one replica is past a disk failure threshold the existing token-level protection can reject writes for the ranges it owns while other writes continue. That can leave an application able to write only part of a keyspace depending on which replicas own each partition. CASSANDRA-21024 adds an option to stop writes across every node replicating the keyspace when any participating node crosses that threshold.
Choosing keyspace-wide rejection makes the failure more consistent for the application but also stops writes that could otherwise succeed. Token-level protection preserves writes to unaffected replica sets at the cost of partial keyspace availability. Test how the application handles each policy because you’ll still have a disk-capacity problem to resolve whichever one you choose.
The driver-version guardrail in CASSANDRA-21146 checks the driver type and version reported when a client connects. It stays disabled until configured so you can first try it against the clients in development or pre-production before using it during an upgrade.
Separate warning and rejection maps let you choose a minimum version for each known driver type. You can also configure unknown and unset entries for clients that don’t supply an identity. Check those rules against an inventory of the real clients so an older service or nonstandard driver doesn’t unexpectedly lose its connection.
A write to a large partition can add to the work required for future reads and compaction long before disk use looks unusual. CASSANDRA-17258 uses node-local top-partition information to warn clients when a target partition already exceeds a tracked size or tombstone threshold. It ships with write_thresholds_enabled: false and unset thresholds so warnings need to be configured before they can help identify those writes. Repair and streaming also have more data to process as the partition grows which is worth including in recovery tests.
| Guardrail | What Cassandra can do | What it does not do |
|---|---|---|
| Disk usage | Warn or reject according to configured thresholds; optionally fail keyspace writes across participating replicas. | Add disk, choose retention, rebalance data, or recover space already consumed by an unsafe table. |
| Driver version | Warn or reject connections that report a configured driver type below the minimum version. | Test an application migration or prove that every service reports its version accurately. |
| Large-partition write | Warn a client when it adds data to a partition already tracked as large or tombstone-heavy. | Redesign the partition key, reduce cardinality, or remove existing oversized partitions. |
| Repair disk protection | Decline repair preparation or streaming under unsafe free-space or compaction conditions. | Create the capacity required to catch up on the repair backlog. |
Set monitoring alerts early enough to give you time to respond before the database reaches a rejection threshold. A growing keyspace should prompt capacity work while writes can still continue without hitting the disk guardrail.
Slow queries, profiling, and metrics
With CASSANDRA-13001 you can query slow-query records through a virtual-table appender instead of relying solely on a debug log destination. A monitoring system can retain those events with the surrounding metrics and logs so the record remains available when you investigate later.
When CPU usage rises and you need to see where the time is going CASSANDRA-20854 exposes async-profiler through JMX and nodetool profile. That provides a profiling route through the normal operational interface which you can test before an incident occurs. Profiling still has a cost so choose and test its use carefully on a busy production node.
The new table and virtual-table metrics let you inspect total rows read and mutated alongside rows-mutated-per-write histograms. You can also examine prepared-statement cache activity and hints with more timer percentile information available for latency checks. Uncaught exceptions and administrative history add records that can be compared with those measurements during an investigation.
| Symptom | Cassandra evidence to correlate | Questions to answer |
|---|---|---|
| Read latency increases | Slow-query records, rows read, SSTables per read, tombstones scanned, read timeouts, device latency, and compaction activity. | Is the request reading more data, touching more SSTables, waiting on disk, or competing with maintenance work? |
| Write latency increases | Rows mutated per write, commitlog and memtable metrics, flush and compaction state, disk queueing, and client errors. | Did mutation size change, did flushing fall behind, or is the storage path saturated? |
| Node CPU is high | Async profile, executor activity, allocation rate, GC information, query and repair rate, and recent configuration events. | Which code path is consuming the time and is it driven by application traffic, repair, compaction, or a regression? |
| A production change precedes drift | nodetool history, configuration history, schema events, topology events, and the before-and-after workload profile. | What changed, when did it propagate, and does the timing align with the observed behaviour? |
When an alert reports a latency increase you’ll want to find the affected table and check what requests were doing there. Comparing that activity with storage work and recent operational changes gives you more to investigate than a node-level CPU or heap graph alone. AxonOps keeps the metrics with logs and events alongside configuration and topology history so you can follow that comparison over time. Repair-control decisions are recorded with that history while Adaptive Repair uses live telemetry to respond to the conditions currently on the cluster.
Cassandra 5.0 and 6.0 validation
Run the 5.0 and 6.0 checks under the maintenance and application workload your production cluster already handles. The table below gives you specific behaviours to compare when repair competes with requests or a guardrail starts rejecting work.
| Test | Cassandra 5.0 baseline | Cassandra 6.0 comparison | Measurements |
|---|---|---|---|
| Repair scheduling | Existing external schedule and its repair history. | Configure the scheduler with conservative assignment, concurrency, and retry settings. | Repair freshness by range and table, bytes repaired, failed and retried sessions, compaction backlog, disk headroom, read/write latency, and history consistency. |
| Disk-pressure response | Existing alerts and token-level write rejection behaviour. | Test configured warning and failure thresholds, including the keyspace-wide option, in a controlled environment. | Per-node free space, rejected writes by token and keyspace, application error handling, replica availability, and recovery after capacity is restored. |
| Driver compatibility | Inventory drivers outside the database. | Warn in pre-production first, then test failure handling with an obsolete declared driver version. | Connection warnings or failures, service rollout readiness, reporting accuracy, and no unplanned application outage. |
| Slow-query analysis | Logs, tracing, JMX, and external profiling workflow. | Add virtual-table slow-query records and a controlled async profile to the incident path. | Query identifiers, table and coordinator context, profile overhead, CPU and allocation evidence, and retained incident history. |
| Large-partition protection | Existing top-partition metrics and manual review. | Test client warnings while a known large partition receives writes. | Warning delivery, partition size and tombstone evolution, read cost, compaction impact, and follow-up schema remediation. |
Contributors
Jaydeepkumar Chovatia proposed and led the Unified Repair Solution and Kristijonas Zalys proposed the still-open changes to concurrent auto-repair history mutations. Isaac Reath developed the keyspace-wide disk-usage guardrail while Brad Schoening raised the driver-version request with Stefan Miklosovic assigned to it. David Capwell raised the large-partition client-warning work and Minal Kyada implemented the change. Jon Haddad proposed the slow-query virtual-table appender while Yaman Ziadeh and Bernardo Botella Corbi brought async profiling to Cassandra’s JMX and nodetool interfaces.
I’m grateful to those contributors and to everyone testing how these features behave when repairs fail or cluster state changes unexpectedly. Reviewing the design and investigating regressions takes time alongside maintaining CI and preparing documentation and releases. Reports from people running Cassandra help that work reach cases that would otherwise be difficult to reproduce.
Series
- Part 1: Notes from Using Cassandra Since 2008
- Part 2: Accord Transactions
- Part 3: Performance Optimisations
- Part 5: Zstd Dictionary Compression
- Part 6: Transactional Cluster Metadata and CMS
- Part 7: Cursor Compaction and SSTable Writes
- Part 8: Storage-Attached Indexing and Schema Constraints
- Part 9: JDK 21 and Generational ZGC
- Part 10: Upgrade and Production Validation
Sources
- Apache Cassandra 6.0 CHANGES.txt
- Apache Cassandra 6.0 cassandra.yaml
- CEP-37: Apache Cassandra Unified Repair Solution
- CASSANDRA-19918: Apache Cassandra Unified Repair Solution
- CASSANDRA-20996: auto-repair history consistency
- CASSANDRA-21024: keyspace-wide disk usage guardrail
- CASSANDRA-21146: client driver version guardrail
- CASSANDRA-17258: warnings for writes to large partitions
- CASSANDRA-13001: slow query virtual table
- CASSANDRA-20854: low-overhead async profiling