Datadog Cassandra Metric Gaps in the Default Integration
The Cassandra Monitoring Tools Comparison 2026 looked at Datadog alongside the tools built specifically for Cassandra. Here we’ll take a table with slow reads and work through what you can check in Datadog before looking at the gaps elsewhere in the cluster.
Datadog gives you local request latency and storage use together with some thread-pool and JVM activity to start that investigation. You’ll have less to work with when checking how replicas recover from an outage because repair progress and hint delivery are missing from the default collection.
Cassandra already exposes node-level metrics before you create a user table and every table adds its own set within its keyspace. With 60 user tables you’re looking at more than 6,900 table metric names on each node before counting individual attributes and percentiles. Keyspace rollups and per-peer metrics add more measurements as you start looking beyond each table.
| Figure | What it shows |
|---|---|
| 72 | Distinct cassandra.* metric names in Datadog's published Cassandra catalogue. |
| ~7 | Cassandra JMX metric families addressed by the default mapping, out of roughly 30 families Cassandra 4.1 and 5.0 publish. |
| 115+ | Metric names Cassandra 5.0 exposes for every table, before attributes, percentiles, or dimensions are selected. |
| 6,900+ | Table metric names Cassandra publishes on one node with 60 user tables, before attributes, percentiles, keyspace rollups, or per-peer metrics. |
| 350 | Datadog's documented collection limit per Cassandra check instance. |
The comparison below uses Datadog’s metrics.yaml and metric catalogue as read on 3 September 2026 against the Apache Cassandra 4.1 and 5.0 metric references. Check the mapping shipped with your installed Agent when following along because it can change between integration releases.
Datadog Default Metrics
You’re looking at a table whose reads have slowed down over the past hour and want to find out what’s changed. Start by comparing its read latency with its request rate to see whether the slowdown coincides with more traffic. If traffic is about the same then it’s worth checking whether each read is touching more SSTables or scanning more tombstones than before. Datadog already collects those measurements so we can follow that part of the investigation without adding anything to the configuration.
The default configuration includes p75 and p95 for ReadLatency and WriteLatency with p99 also available for these table timers. The SSTable and tombstone measurements come from SSTablesPerReadHistogram and TombstoneScannedHistogram while MaxPartitionSize and MeanPartitionSize let you check partition sizes for the same table.
You can compare the table’s disk use and SSTable count with pending compactions and flush activity to check what else changed during that hour. Key-cache hit rate and selected table cache and compression measurements give you further details to examine while system keyspaces are excluded from table-level collection.
Beyond the table you can look for queued work in selected thread pools alongside dropped-message rates and storage exceptions. Commit-log size and pending tasks are available too while host metrics and garbage-collection counters let you compare those changes with resource use. Cassandra’s failure-detector counts show each node’s view of membership and the separate cassandra_nodetool integration adds node state and ownership gauges similar to those reported by nodetool status.
Datadog Metric Gaps
For each question the tables below show which Cassandra JMX metrics you’d want to examine and how much Datadog collects by default.
Repair, hints, streaming, and read repair
| What you need to know | Default status | Cassandra metrics missing from the mapping |
|---|---|---|
| Is data repaired, pending repair, or unrepaired? | Absent | Table PercentRepaired, BytesRepaired, BytesUnrepaired, BytesPendingRepair |
| Are repair jobs progressing or failing? | Absent | RepairJobsStarted, RepairJobsCompleted, RepairTime, RepairPrepareTime, RepairSyncTime, ValidationTime, AnticompactionTime, BytesValidated, PartitionsValidated; Storage:RepairExceptions; Repair retry and preview-failure metrics |
| Are hints accumulating or being delivered? | Absent | Storage:TotalHints, Storage:TotalHintsInProgress; HintsService:HintsSucceeded, HintsFailed, HintsTimedOut, and delay metrics |
| Is bootstrap, rebuild, decommission, or repair streaming underway? | Absent | Streaming:IncomingBytes, OutgoingBytes per peer; TotalIncomingBytes, TotalOutgoingBytes, TotalOutgoingRepairBytes, IncomingProcessTime |
| Are repair-side thread pools saturated? | Absent | AntiEntropyStage, ValidationExecutor, CompactionExecutor, Repair-Task, RepairJobTask, StreamReceiveTask, PaxosRepairStage |
| Is read repair firing or timing out? | Absent | ReadRepair:Attempted, RepairedBlocking, RepairedBackground, RepairTimedOut, SpeculatedRead, SpeculatedWrite; table read-repair request and inconsistency metrics |
Request failures and latency
| What you need to know | Default status | Cassandra metrics missing from the mapping |
|---|---|---|
| Are requests failing or unavailable? | Absent | ClientRequest:Failures, Unavailables across all scopes; Cassandra 5.0 TombstoneAborts and ReadSizeAborts |
| Which consistency level is degraded? | Absent | ClientRequest read and write scopes per consistency level, including LOCAL_QUORUM, EACH_QUORUM, and ONE |
| Is coordinator latency different from replica-local latency? | Absent | Table CoordinatorReadLatency, CoordinatorWriteLatency, CoordinatorScanLatency |
| Is the tail moving beyond p99? | Partial | Max, 999thPercentile, Mean, and total-latency counters; p99 is collected only for table read and write latency |
| Is LWT contention increasing? | Partial | ContentionHistogram, ConditionNotMet, UnfinishedCommit, UnknownResult, and Paxos:LinearizabilityViolations |
| Are materialized-view writes healthy? | Partial | ViewReplicasAttempted, ViewReplicasSuccess, ViewPendingMutations, and view write timeout and failure metrics |
| Are local read and write request rates available? | Covered | OneMinuteRate is collected from the listed latency MBeans |
Compaction, cache, memtable, and commit log detail
| What you need to know | Default status | Cassandra metrics missing from the mapping |
|---|---|---|
| Is compaction keeping up? | Partial | Compaction:PendingTasks, PendingTasksByTableName, CompletedTasks, TotalCompactionsCompleted, BytesCompacted, CompactionsAborted, CompactionsReduced, SSTablesDroppedFromCompaction; the compaction executor pool |
| Is the chunk cache helping the read path? | Absent | Cache scoped to ChunkCache: Hits, Misses, Requests, HitRate, Capacity, Size, Entries, MissLatency |
| Are key, row, and counter caches correctly sized? | Partial | Cache capacity, size, entries, requests, and 1/5/15-minute hit rates; the complete CounterCache family |
| Is memtable pressure building? | Partial | MemtableOnHeapDataSize, MemtableOffHeapDataSize, MemtableLiveDataSize, MemtableSwitchCount, MemtableColumnsCount, AllMemtables*; MemtablePool:BlockedOnAllocation, PendingFlushTasks |
| Is the commit log causing stalls? | Partial | WaitingOnCommit, WaitingOnSegmentAllocation, CompletedTasks, OverSizedMutations |
| Do SSTable layout and compaction state point to a problem? | Partial | EstimatedPartitionCount, UnleveledSSTables, MinPartitionSize, OldVersionSSTableCount, MaxSSTableSize, MaxSSTableDuration, EstimatedPartitionSizeHistogram, SSTableCountPerLevel |
| Is off-heap usage rising through filters and indexes? | Partial | BloomFilterFalsePositives, BloomFilterDiskSpaceUsed, BloomFilterOffHeapMemoryUsed, IndexSummaryOffHeapMemoryUsed, CompressionMetadataOffHeapMemoryUsed |
| Is speculative retry doing useful work or adding load? | Absent | Table SpeculativeRetries, SpeculativeFailedRetries, SpeculativeInsufficientReplicas, AdditionalWrites |
| Is direct-memory or networking buffer pressure visible? | Absent | BufferPool size, used size, capacity, hits, misses, and overflow size for chunk-cache and networking pools |
Internode communication
| What you need to know | Default status | Cassandra metrics missing from the mapping |
|---|---|---|
| Are internode queues backing up, and towards which peer? | Absent | Per-peer Connection metrics for large, small, and urgent message pending tasks and bytes, completed tasks and bytes, drops due to timeout, overload, and error, plus timeouts |
| Is inbound messaging seeing corruption, throttling, or expiry? | Absent | InboundConnection:CorruptFramesRecovered, CorruptFramesUnrecovered, ErrorBytes, ErrorCount, ExpiredBytes, ExpiredCount, ThrottledCount, ThrottledNanos, ProcessedBytes, ScheduledBytes |
| Is cross-node or cross-datacentre latency increasing? | Absent | Messaging:CrossNodeLatency, per-datacentre latency, and per-verb wait latency |
| What caused dropped messages? | Partial | Drop Count, internal and cross-node dropped latency; the default keeps only the one-minute rate |
| Is gossip healthy beyond up/down member counts? | Partial | Failure-detector phi values and simple states, Gossiper state, and the GossipStage pool |
| Is the dynamic snitch penalising replicas? | Absent | DynamicEndpointSnitch:Scores and Severity |
Client, CQL, and workload behaviour
| What you need to know | Default status | Cassandra metrics missing from the mapping |
|---|---|---|
| How many native clients are connected, and under which protocol or user? | Absent | Client:connectedNativeClients, connectedNativeClientsByUser, clientsByProtocolVersion, connections |
| Are authentication failures or protocol errors increasing? | Absent | Client:AuthSuccess, AuthFailure, ProtocolException, UnknownException |
| Is native transport saturated or shedding requests? | Absent | Transport ThreadPools metrics for Native-Transport-Requests; Client:PausedConnections, RequestDiscarded, TimedOutBeforeProcessing, Queued |
| Is the prepared-statement cache thrashing? | Absent | CQL:PreparedStatementsCount, PreparedStatementsEvicted, PreparedStatementsExecuted, RegularStatementsExecuted, PreparedStatementsRatio |
| What is the client request and response volume? | Absent | ClientMessageSize bytes sent and received; ClientRequestSize rows and columns read and written |
| Are oversized batches appearing? | Absent | Batch:PartitionsPerLoggedBatch, PartitionsPerUnloggedBatch, PartitionsPerCounterBatch |
| Are query guardrails warning or aborting work? | Partial | Tombstone, read-size, row-index-size, and live-scanned warning and abort metrics; the default has tombstone-scan percentiles only |
JVM and node state
| What you need to know | Default status | Detail |
|---|---|---|
| Are heap, file descriptors, and thread counts available? | Covered | Datadog collects standard JVM memory, file-descriptor, thread, and buffer metrics |
| Is G1 major-GC activity reported correctly? | Requires verification | The configuration names G1 Mixed Generation; standard HotSpot exposes G1 Old Generation, so test the reported counter on the JVM in use |
| Can pause distribution, allocation rate, or safepoint time be diagnosed? | Absent | The default has cumulative GC count and time, not per-pause duration, allocation rate, promotion, evacuation failure, or safepoint data |
| Are keyspace-level Cassandra rollups available? | Absent | The Keyspace metric family, including WriteFailedIdealCL and IdealCLWriteLatency |
| Can system keyspace and node lifecycle state be inspected? | Absent | System keyspaces are excluded from table collection; StorageService operation mode, bootstrap state, joining/leaving/moving nodes, schema-version state, and related data are not collected |
| Are Storage-Attached Indexes visible in Cassandra 5.0? | Absent | StorageAttachedIndex query timeouts, builds in progress, disk usage, query latency, and memtable index flush errors |
| Are host CPU, disk, and network available? | Covered | Datadog host metrics provide this independently of the Cassandra metric mapping |
Missing Anti-Entropy Metrics
Take a replica that’s come back after an outage and is accepting requests again. You’ll still want to check whether it’s receiving the writes it missed while it was unavailable because a node marked as up doesn’t tell you how far that work has progressed.
The hint-delivery metrics let you follow HintsSucceeded alongside delivery failures and timeouts on the nodes sending hints. If disk or network use rose while data was being streamed then the per-peer byte counters would help you check which nodes were involved.
If a repair job failed you’ll want to find out which token ranges it covered before it stopped. Its completion history helps you identify unfinished work so you can plan the next run around what still needs repairing and the load already on the cluster.
Request Failure and Unavailable Metrics
An application can report unavailable errors while its read and write timeout counters stay much the same. Checking Timeouts alone won’t explain those errors because Datadog leaves Failures and Unavailables out of its default ClientRequest collection.
When you see a timeout the coordinator has waited beyond the request deadline without enough responses to satisfy the requested consistency level. An unavailable exception tells you the coordinator could already see that too few replicas were available to meet that consistency level. A request failure means an error occurred during processing and you’ll need to check the replicas involved to see what went wrong.
If LOCAL_QUORUM requests are failing in one datacentre while requests at ONE succeed then you’ll want to compare the two separately. Cassandra exposes request metrics for each consistency level through JMX but the Datadog configuration needs to collect those scopes and keep the consistency level in their tags.
Missing Latency Metrics
Now suppose the application is still reporting slow reads but the replicas’ local read latency hasn’t changed. A request might be waiting in an outgoing queue on the coordinator before a replica gets the chance to process it. Looking at that peer’s pending tasks and bytes would help you check for a backlog during the slowdown but those connection metrics aren’t in the default collection.
If most reads complete normally while a small proportion take much longer then you’ll want to inspect the upper percentiles for the affected requests. Datadog collects p99 for table reads and writes but several other latency families stop at p95 and the default collection leaves out p99.9 and maximum latency. Table-level coordinator latency and total-latency counters are also missing so there’s less information to compare with the application’s measurements.
| Area | Metrics to compare | Default coverage |
|---|---|---|
| Compaction | Queue depth, completed tasks, bytes compacted, aborted work, and the compaction executor show whether the storage engine is falling behind. | Partial: per-table pending compactions and bytes written only |
| Write path | Memtable memory, allocation blocking, commit-log waits, segment allocation, and oversized mutations separate a full write path from a slow coordinator. | Partial: pending flushes and total commit-log size only |
| Read path | Chunk-cache activity, filter effectiveness, index-summary memory, and compression metadata help explain why reads are now doing more work. | Partial: key-cache hit rate and selected table metrics only |
| Replica path | Per-peer queues, message timeouts, cross-node latency, and replica-local metrics test whether the coordinator is waiting elsewhere. | Absent: not collected |
You might also find that pending compactions increased during the same hour and want to check whether compaction was keeping up. Comparing the queue with completed work and bytes compacted would help you see whether it recovered or continued growing throughout the slowdown. A pending-task chart on its own leaves you with very little to go on about the work being completed during that time.
Missing Internode and Client Metrics
An application deployment might leave the request rate much the same while changing how clients connect and submit queries. You’d want to check whether connections had increased and whether requests were accumulating in transport queues before Cassandra could process them.
Repeated prepared-statement evictions give you another reason to examine client behaviour even when traffic hasn’t increased. The CQL counters would let you check when those evictions occurred and compare them with CPU usage over the same period.
| Area | Metrics to compare | Default coverage |
|---|---|---|
| Peer queues | Large, small, and urgent connection queues, bytes in flight, and drops by timeout, overload, or error identify the affected peer. | Absent: not collected |
| Network transit | Cross-node, per-datacentre, and per-verb wait latency separates a local processing issue from an internode path problem. | Absent: not collected |
| Native transport | Connection counts, authentication failures, protocol errors, transport queues, discarded requests, and pre-processing timeouts describe client pressure. | Absent: not collected |
| CQL behaviour | Prepared-statement evictions, executions, ratios, batch-size histograms, guardrail warnings, and SAI metrics expose query patterns that a basic dashboard cannot. | Absent: not collected |
For a query using SAI you’ll also want to check index-specific query timeouts and whether any index builds are still in progress. The SAI metrics listed above let you examine those details separately from general read activity on the node.
You can already check the tombstone-scan percentiles in Datadog but you’ll need the warning and abort counters to see whether reads crossed the configured limits. A higher scan percentile by itself won’t tell you how many queries were warned about or stopped.
Datadog Metric Limit
Before adding those measurements you’ll need to account for Datadog’s documented limit of 350 metrics per instance for the Cassandra check.
Consider read latency on a cluster with dozens of tables and several percentiles selected for each one. Each table produces separate series on each node before you add the per-peer measurements or split request metrics by consistency level. Cassandra 5.0 exposes more than 115 metric names per table so including more of the measurements used in these checks can quickly take you past the default limit.
Once the data is arriving you’ll need dashboards and alerts that make it useful during an investigation. The runbooks also need to reflect those additions so whoever is handling the incident knows which measurements are available.
The example below selects a few node-level measurements to try before adding metrics for every table or peer.
jmx_metrics:
# Error outcomes in addition to latency and timeouts
- include:
domain: org.apache.cassandra.metrics
type: ClientRequest
name: [Failures, Unavailables]
attribute: [Count, OneMinuteRate]
# Hints and compaction backlog
- include:
domain: org.apache.cassandra.metrics
type: Storage
name: [TotalHints, TotalHintsInProgress]
- include:
domain: org.apache.cassandra.metrics
type: Compaction
name: [PendingTasks, CompletedTasks, BytesCompacted]
# Streaming and native transport pressure
- include:
domain: org.apache.cassandra.metrics
type: Streaming
name: [TotalIncomingBytes, TotalOutgoingBytes]
attribute: [Count]
- include:
domain: org.apache.cassandra.metrics
type: ThreadPools
path: transport
name: [ActiveTasks, PendingTasks, CurrentlyBlockedTasks]
Try this with the Cassandra and Agent versions you’re running and check the resulting metric count and tags before deploying it across the cluster. When you add per-table repair state or per-peer queues you’ll also need room in the collection limit and dashboards that keep those tables and peers identifiable.
Check the garbage-collector names against the JVM on your nodes too because the Datadog configuration includes G1 Mixed Generation for a major collector. Standard HotSpot exposes G1 Young Generation and G1 Old Generation so a name mismatch can leave the major-GC counter empty even when collection is taking place.
Custom Metric Costs
Adding the missing metrics through a custom check can also increase your bill once you’ve used the account’s included custom metric allowance. Datadog’s integration documentation excludes metrics from accepted integrations from that allowance while counting custom-check metrics and unsupported JMX integrations as custom metrics. Before estimating the cost of additional Cassandra metrics you’ll need to confirm which ones Datadog will count as custom metrics.
At the published USD rate checked on 11 September 2026 you pay $5 per 100 indexed custom metrics per month after using the included allowance. The billing rules include 100 custom metrics per host on Pro or 200 on Enterprise and pool that allowance across your account. Datadog counts unique combinations of metric names and tag values including the host tag each hour and averages those hourly counts over the month.
Say you’re monitoring ten Cassandra nodes on Pro and each sends 1000 additional series that Datadog classifies as custom metrics. If those are your only hosts and you haven’t used any of the allowance elsewhere then 1000 of the 10000 series would be included. Keeping the remaining 9000 indexed throughout the month would add $450 per month in additional metric charges to your bill before infrastructure fees. Your contract and other custom metric usage can change the total and this example assumes you haven’t configured Metrics without Limits.
If you configure tags through Metrics without Limits you’ll also pay a separate ingestion charge of $0.10 per 100 ingested custom metrics above the ingestion allowance under the same billing rules.
AxonOps Cassandra Monitoring
Trying to explain yesterday’s slowdown is much harder when the measurements you need were never collected. A latency chart can show when the problem started while leaving you unable to check what compaction or communication between replicas was doing at the time.
AxonOps collects tens of thousands of Cassandra metrics from each node and organises dashboards around the Cassandra operations those measurements describe. If you’re investigating slow reads on one table you can examine its request and storage activity before looking at compaction on the nodes involved.
Repair and backup have their own views of progress and history so you can see what was running during the slowdown. Logs and service checks are available in the same platform with configuration and nodetool activity alongside Cassandra-specific alerts.
AxonOps AI is trained on Cassandra source code and documentation as well as Cassandra operational best practices. When you’re trying to explain a latency change it can examine metrics alongside logs and events while taking the cluster’s configuration and operational history into account.
The Apache Cassandra Monitoring page shows how this monitoring is organised and the Cassandra Monitoring Tools Comparison 2026 covers the other tools considered alongside AxonOps.