Add observability metrics for AI audit events

What

Adds end-to-end observability for the ai_audit_events system (consumer side), across all environments. Metrics only — no retention or schema changes.

Prometheus (GitLab.com):

  • gitlab_ai_audit_events_ingested_total{event_name} — events accepted for storage (counted pre-persistence)
  • gitlab_ai_audit_events_dropped_total{reason} — malformed / unknown_event
  • gitlab_ai_audit_events_stored_total{store} — events durably written to PostgreSQL (store: :postgresql)
  • gitlab_ai_audit_events_buffered_total — events enqueued to the ClickHouse write buffer (Redis)
  • gitlab_ai_audit_events_pg_fallback_total — events stored via the PostgreSQL fallback instead of ClickHouse
  • gitlab_ai_audit_events_streaming_total{streamable,status} on the streaming worker
  • gitlab_clickhouse_write_buffer_pending{table} gauge (generic — covers all ClickHouse write-buffer tables; set before the enabled? guard so the series never freezes)

Service Ping (all environments → data warehouse):

  • counts.count_ai_audit_events (28d + all) — PostgreSQL DatabaseMetric (timestamp_column :created_at, so the 28d frame prunes monthly partitions). Reports ~0 where ClickHouse is the store.
  • counts.count_ai_audit_events_clickhouse (28d + all) — separate GenericMetric gated on globally_enabled_for_analytics?, querying ClickHouse count() inline (no clickhouse data source exists in Service Ping; data_source: database).
  • counts.count_total_ingest_ai_audit_events_batch_monthly (28d) — internal_events count of ingest batches.
  • settings.ai_audit_events_streaming_enabled

Internal Event (all environments → Snowplow/Snowflake):

  • ingest_ai_audit_events_batch — value (batch size) and property (storage backend: clickhouse / postgresql), for cross-environment throughput visibility. Not fired for empty batches.

Why

ai_audit_events is routed to ClickHouse on GitLab.com and PostgreSQL on self-managed/Dedicated (ClickHouse is optional on self-managed). Both grow unbounded today, and rows are large (prompt content and metadata). Before committing to partitioning, sharding, truncation, or a retention policy we need to measure size/throughput and prove coverage in every environment. Service Ping (counts) + Internal Events (throughput) are the only signals that reach the warehouse from self-managed/Dedicated; Prometheus is the GitLab.com high-fidelity supplement. The Rails-side ingested/stored/buffered counters plus the producer-side captured counter (companion ai-gateway MR) form a captured ≈ ingested ≈ stored/buffered reconciliation that proves no loss (divergence is the loss signal, which is why ingested is counted pre-persistence and kept separate from stored/buffered).

Table size & growth observability

Row counts alone under-describe a table with large rows that grows exponentially. Size/growth is observed via existing infrastructure (no new metric needed here):

  • PostgreSQL: Gitlab::Database::PostgresTableSize / postgres_table_sizes (total/table/index/toast bytes, small/medium/large/over_limit, 25 GB alert) and the pg_total_relation_size_bytes{relname="ai_audit_events"} Prometheus series. ai_audit_events is already tracked there.
  • ClickHouse: system.parts (sum(bytes_on_disk), sum(rows) per partition) plus the ClickHouse Prometheus endpoint.
  • Per-partition size/rows and average row size (total_bytes / row_count) are the inputs to the partition/shard/TTL/truncation decisions — captured as the dashboard follow-up below.

Testing

  • Ingest service, process-events worker, streaming worker, write_buffer, and instrumentation metric specs all pass.
  • Metric and internal-event definition validation specs pass.
  • Note: the :click_house-tagged dump_write_buffer_worker examples may fail locally with ClickHouse MEMORY_LIMIT_EXCEEDED (Code 241) — environmental, not a code issue; the gauge logic is also covered by the Redis-level write_buffer tests. They run in CI.

Reviewer notes

  • count_ai_audit_events is now a PostgreSQL DatabaseMetric (SQL generated on GitLab.com and executed later by the Data team), and the ClickHouse count is a separate count_ai_audit_events_clickhouse GenericMetric. This keeps values consistent regardless of whether an instance toggles ClickHouse analytics. Open question for analytics-instrumentation: keep the CH GenericMetric, or drop it if the Data team counts ClickHouse directly? data_source: database is used because Service Ping has no clickhouse source.
  • The Service Ping all-frame PostgreSQL count queries ai_audit_events (potentially large) — database review requested.
  • gitlab_clickhouse_write_buffer_pending lives in the shared ClickHouse::DumpWriteBufferWorker (feature_category: value_stream_management) — please loop in VSM/ClickHouse reviewers.
  • No migration. No changelog: internal observability, no user-facing change.

Out of scope (follow-ups)

  • Retention/TTL + payload truncation (held pending the compliance decision on the retention window).
  • Partitioning/sharding changes for unbounded growth.
  • Promoting buffered_total to a true CH-persisted count (via DumpWriteBufferWorker inserted_rows).
  • Grafana dashboard for per-partition size/rows and average row size (PG + ClickHouse) and alert thresholds for pg_fallback_total / write_buffer_pending (runbooks repo).

Database:

References

Edited by Andrew Jung

Merge request reports

Loading
Loading