Add NATS audit streaming observability metrics

What does this MR do and why?

Adds two Rails-side metrics to observe the NATS audit event streaming pipeline

  • gitlab_audit_event_streaming_nats_publish_duration_seconds (histogram): wall time the synchronous publish blocks the calling Puma/Sidekiq thread. Slow NATS directly adds latency to the request or worker that generated the audit event. Recorded from EnqueueService in an ensure block so it captures every path (ack, fallback, or timeout). Sub-second buckets with a tail bounded by the publish timeout plus connect retries.
  • gitlab_audit_event_streaming_nats_consumer_unreachable_total (counter): consumer drains aborted because NATS was unreachable. Producer-side unreachability already surfaces as publish fallbacks, but the consumer side is otherwise invisible when the Sidekiq fleet is partitioned off from NATS. The counter is pre-created at zero here. The call site will be wired in a follow-up once the consumer worker is on master.

Related: #604455 Epic: &17582

References

Screenshots or screen recordings

Before After

How to set up and validate locally

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Merge request reports

Loading
Loading