All articles

Kafka Monitoring Metrics

Your producers are taking longer to receive acknowledgements and consumer lag is climbing even though broker CPU and incoming traffic look much the same as yesterday. Before adding capacity we can look at where requests spend their time and follow the delay through the affected brokers and applications.

A producer might be waiting for replicas to catch up while a consumer falls behind because each record triggers a slow database write. We need measurements from the brokers and the applications to work out which part of that process is slowing down.

We checked the metrics and commands against a local Kafka 4.1.0 cluster with four brokers and three KRaft controllers using definitions from the Apache 4.1 documentation and 4.1.0 source. The captured results below come from those checks while the worked 20 ms request is illustrative. Availability can vary with the deployed Kafka release and client library so the examples need to be matched to your own deployment.

Broker Metrics

Let’s start by comparing the affected brokers so a busy node doesn’t get hidden inside a cluster average. The Apache monitoring reference lists these JMX objects with request-specific measurements where applicable although an exporter may rename them or expose their attributes as separate series.

JMX metricUnit and attributeHow to use it
kafka.server:type=BrokerTopicMetrics,name=BytesInPerSecBytes/s through a rate attribute such as OneMinuteRateCompare incoming traffic across brokers and with previous periods
kafka.server:type=BrokerTopicMetrics,name=BytesOutPerSecBytes/s through the same rate windowLook for extra consumer traffic and replays
kafka.network:type=RequestMetrics,name=TotalTimeMs,request=ProduceMilliseconds through histogram attributes such as Mean and 99thPercentileTrack broker-side produce latency
kafka.network:type=RequestMetrics,name=RequestQueueTimeMs,request=ProduceMilliseconds through histogram attributesIdentify time waiting for request handling
kafka.server:type=ReplicaManager,name=UnderReplicatedPartitionsPartition count through ValueFind partitions with fewer in-sync replicas than their replication factor
kafka.server:type=ReplicaManager,name=UnderMinIsrPartitionCountPartition count through ValueIdentify leader partitions below the configured minimum ISR

The replication gauges in ReplicaManager help us check whether replicas have fallen behind while the partition still has a leader serving requests. Reduced redundancy can leave a partition available so we need to look at both its ISR and leader status before deciding what has failed.

To narrow the traffic comparison to a particular topic we can use the corresponding objects with a topic property. The broker total already includes that traffic so adding the total to its topic series would count the same bytes twice.

Request Latency

If the affected brokers show slower produce requests we can break the time down using the stages recorded in RequestChannel. The following metric names use kafka.network:type=RequestMetrics with request=Produce to separate local processing from time spent waiting afterwards.

Let’s follow a hypothetical produce request that uses acks=all and takes 20 ms on the broker without any quota throttling.

StageMetric nameIllustrative durationNext check
Waiting for a request handlerRequestQueueTimeMs2 msRequest load and handler utilisation
Local processingLocalTimeMs3 msBroker CPU and local I/O
Waiting after local processingRemoteTimeMs12 msFollower replication and the ISR for the affected partitions
Waiting to send the responseResponseQueueTimeMs1 msNetwork-thread availability
Sending the responseResponseSendTimeMs2 msNetwork and client receive behaviour
Total broker request timeTotalTimeMs20 msCompare with the producer’s own request latency

Of those 20 ms the request spends 12 ms waiting after local processing and only 2 ms waiting for a handler. That makes follower replication worth checking before adding handler threads although we still need the follower’s metrics to distinguish disk delays from CPU or network pressure.

We can add these stage durations because they belong to the same hypothetical request but adding the p99 values from separate histograms won’t give the total p99 latency. Each histogram can have different requests at its slow end so use TotalTimeMs for the total distribution and compare the individual stages alongside it.

When quotas apply we should also inspect ThrottleTimeMs to see whether Kafka is deliberately delaying requests. Fetch requests can spend time waiting for more data under their fetch settings so an increase in fetch duration needs to be checked against that configuration.

The local broker exposed Mean and Count alongside 99thPercentile for each of the request stages above. These were histogram samples collected through JMX so they let us check the available attributes without claiming to trace an individual request through every stage.

Replication Problems

For the replication check we’ll use a topic with replication factor three and min.insync.replicas=2 while the producer uses acks=all. These examples have Eligible Leader Replicas disabled because enabling ELR changes the behaviour we need to consider.

Partition stateWhat it means for writesInvestigation
Three replicas in the ISRWrites wait for the in-sync replicasCompare replication delay when acknowledgements slow down
Two replicas in the ISRWrites can still succeed but one replica has fallen out of syncFind the missing replica and check its recovery progress
One replica in the ISRWrites using acks=all cannot meet the configured minimumRestore enough in-sync replicas and check producer errors
No available leaderClients cannot use a healthy leader for that partitionInvestigate broker availability and leader election

With all three replicas in the ISR the producer still waits for the in-sync replicas even though the minimum ISR configuration is set to two. That setting determines the minimum required for these writes without reducing a healthy three-replica ISR to a two-replica acknowledgement rule. A deployment with ELR enabled needs its feature-specific semantics checked separately.

The read-only topic description below shows the leader and ISR so we can identify which replicas to investigate. Replace the bootstrap address and use a client configuration containing the TLS and authentication settings for your cluster.

bin/kafka-topics.sh \
  --bootstrap-server broker.example.com:9092 \
  --command-config client.properties \
  --describe --topic orders

Once we know which replica is missing from the ISR we can check its broker logs alongside its recovery progress. Lowering the minimum ISR changes the durability requirements under which writes may continue but leaves the replica problem unresolved.

For the local check we assigned the topic’s replicas to brokers 1,3,4 while controllers ran on nodes 1 through 3. Stopping broker 4 and then node 3 reduced the ISR without losing the controller majority on nodes 1 and 2.

Captured ISRObserved write result with acks=all
1,3,4Accepted at offset 0
1,3Accepted at offset 1
1Broker returned NOT_ENOUGH_REPLICAS and the producer exhausted its eight-second delivery timeout

The surviving controllers continued to make metadata progress during the final check so we could distinguish the minimum-ISR write failure from a controller quorum failure. These broker shutdowns belong in an isolated test environment and aren’t part of the read-only investigation commands in this guide.

KRaft Monitoring

If metadata operations are also getting stuck we need to look at the KRaft controllers that maintain cluster metadata. Their quorum is separate from topic replication so consumers can still be fetching records while metadata changes are failing to progress.

MetricProcess and unitInterpretation
kafka.controller:type=KafkaController,name=ActiveControllerCountController gauge with Value of zero or oneExpect one active controller across a healthy controller quorum
kafka.server:type=raft-metrics attribute current-stateRaft member state stringIdentify the leader and followers while investigating elections
kafka.server:type=raft-metrics attribute high-watermarkMetadata log offsetCheck progress over time when metadata changes are expected
kafka.server:type=raft-metrics attribute commit-latency-avgMilliseconds on the quorum leaderInvestigate slower metadata log commits

Keep these measurements labelled by controller node so you can compare the leader with its followers using the definitions in QuorumControllerMetrics and KafkaRaftMetrics. If samples stop arriving we first need to check collection because missing data doesn’t tell us that the controller has become inactive.

To inspect the quorum directly we can run the following read-only status check using the metadata quorum tool with the same client configuration.

bin/kafka-metadata-quorum.sh \
  --bootstrap-server broker.example.com:9092 \
  --command-config client.properties \
  describe --status

A metadata offset that stays unchanged while operations are waiting gives us a reason to investigate the quorum. During a quiet period with no metadata changes to commit we wouldn’t expect the offset to keep advancing.

Our three controllers reported ActiveControllerCount values of 0 / 1 / 0 and the following attributes came from the raft metrics on a follower.

current-state=follower
high-watermark=130.0
commit-latency-avg=NaN

The follower’s NaN value isn’t a measurement of zero commit latency so it shouldn’t be converted into a reassuring zero on a chart. For deployments with dedicated controller processes we also need collection from those JVMs because a broker-only installation can’t establish controller coverage.

Producer and Consumer Metrics

If broker timings don’t explain the application’s delay we need measurements from the process running the producer or consumer. A broker JVM doesn’t expose the Java client’s local metrics and other client libraries can have their own names and collection methods.

Java client attributeMetric groupUnitUse
record-queue-time-avgkafka.producer:type=producer-metrics,client-id=...MillisecondsTime batches spend in the producer’s send buffer
request-latency-avgSame producer groupMillisecondsRequest latency seen by the producer
record-retry-rateSame producer groupRetried record sends/sRepeated sends that need an error and timeout investigation
records-lag-maxkafka.consumer:type=consumer-fetch-manager-metrics,client-id=...Offset-distance lagLargest partition lag in the window based on the consumer’s current position
records-consumed-rateSame consumer groupRecords/sFetch consumption rate rather than confirmed downstream completion
fetch-latency-avgSame consumer groupMillisecondsDuration of consumer fetch requests

The producer’s measurements are defined in SenderMetricsRegistry while FetchMetricsRegistry defines the consumer measurements used here. A consumer that has fetched records but hasn’t finished processing them can have a current position well ahead of its committed offset.

For a producer we can compare queue time with the batching delay configured through linger.ms before deciding whether it is waiting unexpectedly. Some time in the send buffer can be intentional because the producer is collecting records into a batch.

The consumer lag guide follows an orders topic through the difference between committed offsets and processing rates. If those records trigger work in another system we also need to measure when that work finishes to understand the delay experienced by users.

Kafka Connect Metrics

For a Kafka Connect sink we can make a similar comparison between records read from Kafka and records delivered to the destination. A task can keep reading successfully while a slow destination leaves more of its work unfinished.

Sink-task attributeUnitInterpretation
sink-record-read-rateRecords/sReads before transformations
sink-record-send-rateRecords/sRecords passed to the sink task after transformations
sink-record-active-countRecordsRead records not yet completely committed or acknowledged by the task

The sink-task attributes use kafka.connect:type=sink-task-metrics,connector=...,task=... and are recorded by WorkerSinkTask. Read and send rates can differ when a transformation filters records so we should compare growing active-record counts with destination latency and task logs before concluding that delivery has slowed.

We checked these attributes with a local Connect 4.1.0 FileStreamSinkConnector that wrote 200 records to a temporary file without any filtering transformations. The connector and task both reported RUNNING while the immediate JMX sample still showed sink-record-active-count=200.0.

The records had reached the file while the task still counted them as active at that sampling point. We didn’t capture the later transition to zero so a single active-count sample wouldn’t justify diagnosing a stuck sink without checking its progress over time.

Alerts and Investigation

Once we understand the source of the delay we can set alerts around how long the application can reasonably wait. A backlog that drains after a scheduled batch job needs a different response from a payment consumer that has stopped processing.

ConditionUseful alert context
Under-replicated partitions persistAffected brokers and topics plus maintenance activity
Produce latency risesRequest stages and producer errors over the same window
Lag grows continuouslyIncoming rate and completed processing rate plus the affected partitions
Controller state changes repeatedlyQuorum membership and metadata operation failures
Disk free space is decliningGrowth rate and room for recovery or reassignment
Telemetry stops arrivingCollection failure distinguished from a genuine zero value

Keep enough history at the original sampling resolution to compare a short slowdown with the period before it began. Use the same time windows when comparing rates and keep broker percentiles separate because averaging them doesn’t produce a cluster-wide percentile.

Kafka Monitoring in AxonOps

In AxonOps we can move between broker and topic views and inspect consumer groups or Connect as the investigation develops. Those views can be checked alongside logs and events to follow the period when the delay began. The screenshot below shows an existing consumer-group view rather than a capture of the local test workload.

AxonOps consumer-group list showing group state, membership and lag

The Kafka monitoring page shows the dashboards and alerting controls available for those checks. Client-side measurements still need collection from the relevant application integration so a broker agent alone shouldn’t be assumed to provide every metric in this guide.

The local exercise verified Kafka behaviour through its APIs and JMX without an AxonOps installation. It therefore doesn’t establish which of these attributes an AxonOps deployment collects or how its dashboards aggregate them.

Our Kafka monitoring and operations overview covers how these views fit into day-to-day Kafka operations. For the delayed acknowledgements we started with we can now compare the broker’s request stages with replication progress and the producer’s own timing before deciding what needs to change.

All articles