Loading
Add NATS audit streaming observability metrics
What does this MR do and why?
Adds two Rails-side metrics to observe the NATS audit event streaming pipeline
gitlab_audit_event_streaming_nats_publish_duration_seconds(histogram): wall time the synchronous publish blocks the calling Puma/Sidekiq thread. Slow NATS directly adds latency to the request or worker that generated the audit event. Recorded fromEnqueueServicein an ensure block so it captures every path (ack, fallback, or timeout). Sub-second buckets with a tail bounded by the publish timeout plus connect retries.gitlab_audit_event_streaming_nats_consumer_unreachable_total(counter): consumer drains aborted because NATS was unreachable. Producer-side unreachability already surfaces as publish fallbacks, but the consumer side is otherwise invisible when the Sidekiq fleet is partitioned off from NATS. The counter is pre-created at zero here. The call site will be wired in a follow-up once the consumer worker is on master.
References
Screenshots or screen recordings
| Before | After |
|---|---|
How to set up and validate locally
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.