All articles

Datadog Cassandra Metric Gaps

Datadog Cassandra Metric Gaps in the Default Integration

The Cassandra Monitoring Tools Comparison 2026 looked at Datadog alongside the tools built specifically for Cassandra. Here we’ll take a table with slow reads and work through what you can check in Datadog before looking at the gaps elsewhere in the cluster.

Datadog gives you local request latency and storage use together with some thread-pool and JVM activity to start that investigation. You’ll have less to work with when checking how replicas recover from an outage because repair progress and hint delivery are missing from the default collection.

Cassandra already exposes node-level metrics before you create a user table and every table adds its own set within its keyspace. With 60 user tables you’re looking at more than 6,900 table metric names on each node before counting individual attributes and percentiles. Keyspace rollups and per-peer metrics add more measurements as you start looking beyond each table.

Figure What it shows
72 Distinct cassandra.* metric names in Datadog's published Cassandra catalogue.
~7 Cassandra JMX metric families addressed by the default mapping, out of roughly 30 families Cassandra 4.1 and 5.0 publish.
115+ Metric names Cassandra 5.0 exposes for every table, before attributes, percentiles, or dimensions are selected.
6,900+ Table metric names Cassandra publishes on one node with 60 user tables, before attributes, percentiles, keyspace rollups, or per-peer metrics.
350 Datadog's documented collection limit per Cassandra check instance.

The comparison below uses Datadog’s metrics.yaml and metric catalogue as read on 3 September 2026 against the Apache Cassandra 4.1 and 5.0 metric references. Check the mapping shipped with your installed Agent when following along because it can change between integration releases.

Datadog Default Metrics

You’re looking at a table whose reads have slowed down over the past hour and want to find out what’s changed. Start by comparing its read latency with its request rate to see whether the slowdown coincides with more traffic. If traffic is about the same then it’s worth checking whether each read is touching more SSTables or scanning more tombstones than before. Datadog already collects those measurements so we can follow that part of the investigation without adding anything to the configuration.

The default configuration includes p75 and p95 for ReadLatency and WriteLatency with p99 also available for these table timers. The SSTable and tombstone measurements come from SSTablesPerReadHistogram and TombstoneScannedHistogram while MaxPartitionSize and MeanPartitionSize let you check partition sizes for the same table.

You can compare the table’s disk use and SSTable count with pending compactions and flush activity to check what else changed during that hour. Key-cache hit rate and selected table cache and compression measurements give you further details to examine while system keyspaces are excluded from table-level collection.

Beyond the table you can look for queued work in selected thread pools alongside dropped-message rates and storage exceptions. Commit-log size and pending tasks are available too while host metrics and garbage-collection counters let you compare those changes with resource use. Cassandra’s failure-detector counts show each node’s view of membership and the separate cassandra_nodetool integration adds node state and ownership gauges similar to those reported by nodetool status.

Datadog Metric Gaps

For each question the tables below show which Cassandra JMX metrics you’d want to examine and how much Datadog collects by default.

Repair, hints, streaming, and read repair

What you need to knowDefault statusCassandra metrics missing from the mapping
Is data repaired, pending repair, or unrepaired?AbsentTable PercentRepaired, BytesRepaired, BytesUnrepaired, BytesPendingRepair
Are repair jobs progressing or failing?AbsentRepairJobsStarted, RepairJobsCompleted, RepairTime, RepairPrepareTime, RepairSyncTime, ValidationTime, AnticompactionTime, BytesValidated, PartitionsValidated; Storage:RepairExceptions; Repair retry and preview-failure metrics
Are hints accumulating or being delivered?AbsentStorage:TotalHints, Storage:TotalHintsInProgress; HintsService:HintsSucceeded, HintsFailed, HintsTimedOut, and delay metrics
Is bootstrap, rebuild, decommission, or repair streaming underway?AbsentStreaming:IncomingBytes, OutgoingBytes per peer; TotalIncomingBytes, TotalOutgoingBytes, TotalOutgoingRepairBytes, IncomingProcessTime
Are repair-side thread pools saturated?AbsentAntiEntropyStage, ValidationExecutor, CompactionExecutor, Repair-Task, RepairJobTask, StreamReceiveTask, PaxosRepairStage
Is read repair firing or timing out?AbsentReadRepair:Attempted, RepairedBlocking, RepairedBackground, RepairTimedOut, SpeculatedRead, SpeculatedWrite; table read-repair request and inconsistency metrics

Request failures and latency

What you need to knowDefault statusCassandra metrics missing from the mapping
Are requests failing or unavailable?AbsentClientRequest:Failures, Unavailables across all scopes; Cassandra 5.0 TombstoneAborts and ReadSizeAborts
Which consistency level is degraded?AbsentClientRequest read and write scopes per consistency level, including LOCAL_QUORUM, EACH_QUORUM, and ONE
Is coordinator latency different from replica-local latency?AbsentTable CoordinatorReadLatency, CoordinatorWriteLatency, CoordinatorScanLatency
Is the tail moving beyond p99?PartialMax, 999thPercentile, Mean, and total-latency counters; p99 is collected only for table read and write latency
Is LWT contention increasing?PartialContentionHistogram, ConditionNotMet, UnfinishedCommit, UnknownResult, and Paxos:LinearizabilityViolations
Are materialized-view writes healthy?PartialViewReplicasAttempted, ViewReplicasSuccess, ViewPendingMutations, and view write timeout and failure metrics
Are local read and write request rates available?CoveredOneMinuteRate is collected from the listed latency MBeans

Compaction, cache, memtable, and commit log detail

What you need to knowDefault statusCassandra metrics missing from the mapping
Is compaction keeping up?PartialCompaction:PendingTasks, PendingTasksByTableName, CompletedTasks, TotalCompactionsCompleted, BytesCompacted, CompactionsAborted, CompactionsReduced, SSTablesDroppedFromCompaction; the compaction executor pool
Is the chunk cache helping the read path?AbsentCache scoped to ChunkCache: Hits, Misses, Requests, HitRate, Capacity, Size, Entries, MissLatency
Are key, row, and counter caches correctly sized?PartialCache capacity, size, entries, requests, and 1/5/15-minute hit rates; the complete CounterCache family
Is memtable pressure building?PartialMemtableOnHeapDataSize, MemtableOffHeapDataSize, MemtableLiveDataSize, MemtableSwitchCount, MemtableColumnsCount, AllMemtables*; MemtablePool:BlockedOnAllocation, PendingFlushTasks
Is the commit log causing stalls?PartialWaitingOnCommit, WaitingOnSegmentAllocation, CompletedTasks, OverSizedMutations
Do SSTable layout and compaction state point to a problem?PartialEstimatedPartitionCount, UnleveledSSTables, MinPartitionSize, OldVersionSSTableCount, MaxSSTableSize, MaxSSTableDuration, EstimatedPartitionSizeHistogram, SSTableCountPerLevel
Is off-heap usage rising through filters and indexes?PartialBloomFilterFalsePositives, BloomFilterDiskSpaceUsed, BloomFilterOffHeapMemoryUsed, IndexSummaryOffHeapMemoryUsed, CompressionMetadataOffHeapMemoryUsed
Is speculative retry doing useful work or adding load?AbsentTable SpeculativeRetries, SpeculativeFailedRetries, SpeculativeInsufficientReplicas, AdditionalWrites
Is direct-memory or networking buffer pressure visible?AbsentBufferPool size, used size, capacity, hits, misses, and overflow size for chunk-cache and networking pools

Internode communication

What you need to knowDefault statusCassandra metrics missing from the mapping
Are internode queues backing up, and towards which peer?AbsentPer-peer Connection metrics for large, small, and urgent message pending tasks and bytes, completed tasks and bytes, drops due to timeout, overload, and error, plus timeouts
Is inbound messaging seeing corruption, throttling, or expiry?AbsentInboundConnection:CorruptFramesRecovered, CorruptFramesUnrecovered, ErrorBytes, ErrorCount, ExpiredBytes, ExpiredCount, ThrottledCount, ThrottledNanos, ProcessedBytes, ScheduledBytes
Is cross-node or cross-datacentre latency increasing?AbsentMessaging:CrossNodeLatency, per-datacentre latency, and per-verb wait latency
What caused dropped messages?PartialDrop Count, internal and cross-node dropped latency; the default keeps only the one-minute rate
Is gossip healthy beyond up/down member counts?PartialFailure-detector phi values and simple states, Gossiper state, and the GossipStage pool
Is the dynamic snitch penalising replicas?AbsentDynamicEndpointSnitch:Scores and Severity

Client, CQL, and workload behaviour

What you need to knowDefault statusCassandra metrics missing from the mapping
How many native clients are connected, and under which protocol or user?AbsentClient:connectedNativeClients, connectedNativeClientsByUser, clientsByProtocolVersion, connections
Are authentication failures or protocol errors increasing?AbsentClient:AuthSuccess, AuthFailure, ProtocolException, UnknownException
Is native transport saturated or shedding requests?AbsentTransport ThreadPools metrics for Native-Transport-Requests; Client:PausedConnections, RequestDiscarded, TimedOutBeforeProcessing, Queued
Is the prepared-statement cache thrashing?AbsentCQL:PreparedStatementsCount, PreparedStatementsEvicted, PreparedStatementsExecuted, RegularStatementsExecuted, PreparedStatementsRatio
What is the client request and response volume?AbsentClientMessageSize bytes sent and received; ClientRequestSize rows and columns read and written
Are oversized batches appearing?AbsentBatch:PartitionsPerLoggedBatch, PartitionsPerUnloggedBatch, PartitionsPerCounterBatch
Are query guardrails warning or aborting work?PartialTombstone, read-size, row-index-size, and live-scanned warning and abort metrics; the default has tombstone-scan percentiles only

JVM and node state

What you need to knowDefault statusDetail
Are heap, file descriptors, and thread counts available?CoveredDatadog collects standard JVM memory, file-descriptor, thread, and buffer metrics
Is G1 major-GC activity reported correctly?Requires verificationThe configuration names G1 Mixed Generation; standard HotSpot exposes G1 Old Generation, so test the reported counter on the JVM in use
Can pause distribution, allocation rate, or safepoint time be diagnosed?AbsentThe default has cumulative GC count and time, not per-pause duration, allocation rate, promotion, evacuation failure, or safepoint data
Are keyspace-level Cassandra rollups available?AbsentThe Keyspace metric family, including WriteFailedIdealCL and IdealCLWriteLatency
Can system keyspace and node lifecycle state be inspected?AbsentSystem keyspaces are excluded from table collection; StorageService operation mode, bootstrap state, joining/leaving/moving nodes, schema-version state, and related data are not collected
Are Storage-Attached Indexes visible in Cassandra 5.0?AbsentStorageAttachedIndex query timeouts, builds in progress, disk usage, query latency, and memtable index flush errors
Are host CPU, disk, and network available?CoveredDatadog host metrics provide this independently of the Cassandra metric mapping

Missing Anti-Entropy Metrics

Take a replica that’s come back after an outage and is accepting requests again. You’ll still want to check whether it’s receiving the writes it missed while it was unavailable because a node marked as up doesn’t tell you how far that work has progressed.

The hint-delivery metrics let you follow HintsSucceeded alongside delivery failures and timeouts on the nodes sending hints. If disk or network use rose while data was being streamed then the per-peer byte counters would help you check which nodes were involved.

If a repair job failed you’ll want to find out which token ranges it covered before it stopped. Its completion history helps you identify unfinished work so you can plan the next run around what still needs repairing and the load already on the cluster.

Request Failure and Unavailable Metrics

An application can report unavailable errors while its read and write timeout counters stay much the same. Checking Timeouts alone won’t explain those errors because Datadog leaves Failures and Unavailables out of its default ClientRequest collection.

When you see a timeout the coordinator has waited beyond the request deadline without enough responses to satisfy the requested consistency level. An unavailable exception tells you the coordinator could already see that too few replicas were available to meet that consistency level. A request failure means an error occurred during processing and you’ll need to check the replicas involved to see what went wrong.

If LOCAL_QUORUM requests are failing in one datacentre while requests at ONE succeed then you’ll want to compare the two separately. Cassandra exposes request metrics for each consistency level through JMX but the Datadog configuration needs to collect those scopes and keep the consistency level in their tags.

Missing Latency Metrics

Now suppose the application is still reporting slow reads but the replicas’ local read latency hasn’t changed. A request might be waiting in an outgoing queue on the coordinator before a replica gets the chance to process it. Looking at that peer’s pending tasks and bytes would help you check for a backlog during the slowdown but those connection metrics aren’t in the default collection.

If most reads complete normally while a small proportion take much longer then you’ll want to inspect the upper percentiles for the affected requests. Datadog collects p99 for table reads and writes but several other latency families stop at p95 and the default collection leaves out p99.9 and maximum latency. Table-level coordinator latency and total-latency counters are also missing so there’s less information to compare with the application’s measurements.

AreaMetrics to compareDefault coverage
CompactionQueue depth, completed tasks, bytes compacted, aborted work, and the compaction executor show whether the storage engine is falling behind.Partial: per-table pending compactions and bytes written only
Write pathMemtable memory, allocation blocking, commit-log waits, segment allocation, and oversized mutations separate a full write path from a slow coordinator.Partial: pending flushes and total commit-log size only
Read pathChunk-cache activity, filter effectiveness, index-summary memory, and compression metadata help explain why reads are now doing more work.Partial: key-cache hit rate and selected table metrics only
Replica pathPer-peer queues, message timeouts, cross-node latency, and replica-local metrics test whether the coordinator is waiting elsewhere.Absent: not collected

You might also find that pending compactions increased during the same hour and want to check whether compaction was keeping up. Comparing the queue with completed work and bytes compacted would help you see whether it recovered or continued growing throughout the slowdown. A pending-task chart on its own leaves you with very little to go on about the work being completed during that time.

Missing Internode and Client Metrics

An application deployment might leave the request rate much the same while changing how clients connect and submit queries. You’d want to check whether connections had increased and whether requests were accumulating in transport queues before Cassandra could process them.

Repeated prepared-statement evictions give you another reason to examine client behaviour even when traffic hasn’t increased. The CQL counters would let you check when those evictions occurred and compare them with CPU usage over the same period.

AreaMetrics to compareDefault coverage
Peer queuesLarge, small, and urgent connection queues, bytes in flight, and drops by timeout, overload, or error identify the affected peer.Absent: not collected
Network transitCross-node, per-datacentre, and per-verb wait latency separates a local processing issue from an internode path problem.Absent: not collected
Native transportConnection counts, authentication failures, protocol errors, transport queues, discarded requests, and pre-processing timeouts describe client pressure.Absent: not collected
CQL behaviourPrepared-statement evictions, executions, ratios, batch-size histograms, guardrail warnings, and SAI metrics expose query patterns that a basic dashboard cannot.Absent: not collected

For a query using SAI you’ll also want to check index-specific query timeouts and whether any index builds are still in progress. The SAI metrics listed above let you examine those details separately from general read activity on the node.

You can already check the tombstone-scan percentiles in Datadog but you’ll need the warning and abort counters to see whether reads crossed the configured limits. A higher scan percentile by itself won’t tell you how many queries were warned about or stopped.

Datadog Metric Limit

Before adding those measurements you’ll need to account for Datadog’s documented limit of 350 metrics per instance for the Cassandra check.

Consider read latency on a cluster with dozens of tables and several percentiles selected for each one. Each table produces separate series on each node before you add the per-peer measurements or split request metrics by consistency level. Cassandra 5.0 exposes more than 115 metric names per table so including more of the measurements used in these checks can quickly take you past the default limit.

Once the data is arriving you’ll need dashboards and alerts that make it useful during an investigation. The runbooks also need to reflect those additions so whoever is handling the incident knows which measurements are available.

The example below selects a few node-level measurements to try before adding metrics for every table or peer.

jmx_metrics:
  # Error outcomes in addition to latency and timeouts
  - include:
      domain: org.apache.cassandra.metrics
      type: ClientRequest
      name: [Failures, Unavailables]
      attribute: [Count, OneMinuteRate]

  # Hints and compaction backlog
  - include:
      domain: org.apache.cassandra.metrics
      type: Storage
      name: [TotalHints, TotalHintsInProgress]
  - include:
      domain: org.apache.cassandra.metrics
      type: Compaction
      name: [PendingTasks, CompletedTasks, BytesCompacted]

  # Streaming and native transport pressure
  - include:
      domain: org.apache.cassandra.metrics
      type: Streaming
      name: [TotalIncomingBytes, TotalOutgoingBytes]
      attribute: [Count]
  - include:
      domain: org.apache.cassandra.metrics
      type: ThreadPools
      path: transport
      name: [ActiveTasks, PendingTasks, CurrentlyBlockedTasks]

Try this with the Cassandra and Agent versions you’re running and check the resulting metric count and tags before deploying it across the cluster. When you add per-table repair state or per-peer queues you’ll also need room in the collection limit and dashboards that keep those tables and peers identifiable.

Check the garbage-collector names against the JVM on your nodes too because the Datadog configuration includes G1 Mixed Generation for a major collector. Standard HotSpot exposes G1 Young Generation and G1 Old Generation so a name mismatch can leave the major-GC counter empty even when collection is taking place.

Custom Metric Costs

Adding the missing metrics through a custom check can also increase your bill once you’ve used the account’s included custom metric allowance. Datadog’s integration documentation excludes metrics from accepted integrations from that allowance while counting custom-check metrics and unsupported JMX integrations as custom metrics. Before estimating the cost of additional Cassandra metrics you’ll need to confirm which ones Datadog will count as custom metrics.

At the published USD rate checked on 11 September 2026 you pay $5 per 100 indexed custom metrics per month after using the included allowance. The billing rules include 100 custom metrics per host on Pro or 200 on Enterprise and pool that allowance across your account. Datadog counts unique combinations of metric names and tag values including the host tag each hour and averages those hourly counts over the month.

Say you’re monitoring ten Cassandra nodes on Pro and each sends 1000 additional series that Datadog classifies as custom metrics. If those are your only hosts and you haven’t used any of the allowance elsewhere then 1000 of the 10000 series would be included. Keeping the remaining 9000 indexed throughout the month would add $450 per month in additional metric charges to your bill before infrastructure fees. Your contract and other custom metric usage can change the total and this example assumes you haven’t configured Metrics without Limits.

If you configure tags through Metrics without Limits you’ll also pay a separate ingestion charge of $0.10 per 100 ingested custom metrics above the ingestion allowance under the same billing rules.

AxonOps Cassandra Monitoring

Trying to explain yesterday’s slowdown is much harder when the measurements you need were never collected. A latency chart can show when the problem started while leaving you unable to check what compaction or communication between replicas was doing at the time.

AxonOps collects tens of thousands of Cassandra metrics from each node and organises dashboards around the Cassandra operations those measurements describe. If you’re investigating slow reads on one table you can examine its request and storage activity before looking at compaction on the nodes involved.

Repair and backup have their own views of progress and history so you can see what was running during the slowdown. Logs and service checks are available in the same platform with configuration and nodetool activity alongside Cassandra-specific alerts.

AxonOps AI is trained on Cassandra source code and documentation as well as Cassandra operational best practices. When you’re trying to explain a latency change it can examine metrics alongside logs and events while taking the cluster’s configuration and operational history into account.

The Apache Cassandra Monitoring page shows how this monitoring is organised and the Cassandra Monitoring Tools Comparison 2026 covers the other tools considered alongside AxonOps.

Sources

  1. Datadog Cassandra integration default metric configuration
  2. Datadog Cassandra integration metric catalogue
  3. Datadog Cassandra integration documentation
  4. Datadog Cassandra nodetool integration documentation
  5. Apache Cassandra metrics reference
  6. Apache Cassandra 5.0 metrics source
All articles