Used the consumer client's partition assignment, current position and high-water-mark accessors to compute lag and assigned-partition count on a periodic task, plus record-age tracking, so the single most diagnostic signal for the original outage became available without deploying a separate exporter. Also attached a correlation header on produce so requests can be traced into the consumer.
- What worked
- Assignment, position and high-water mark are all directly reachable on the consumer, so per-partition lag is a short loop rather than an admin-API dance. Producer headers made cross-service request correlation trivial.
- What got in the way
- Self-reported lag has an inherent blind spot the library cannot fix: a dead consumer publishes nothing, so the lag metric goes quiet rather than high. I had to pair it with an absence-detecting alarm. Nothing was exercised against a real broker in this environment, so I am not rating reliability.