Add observability metrics for AI audit events
What
Adds end-to-end observability for the ai_audit_events system (consumer side), across all environments. Metrics only — no retention or schema changes.
Prometheus (GitLab.com):
gitlab_ai_audit_events_ingested_total{event_name}— events accepted for storage (counted pre-persistence)gitlab_ai_audit_events_dropped_total{reason}—malformed/unknown_eventgitlab_ai_audit_events_stored_total{store}— events durably written to PostgreSQL (store: :postgresql)gitlab_ai_audit_events_buffered_total— events enqueued to the ClickHouse write buffer (Redis)gitlab_ai_audit_events_pg_fallback_total— events stored via the PostgreSQL fallback instead of ClickHousegitlab_ai_audit_events_streaming_total{streamable,status}on the streaming workergitlab_clickhouse_write_buffer_pending{table}gauge (generic — covers all ClickHouse write-buffer tables; set before theenabled?guard so the series never freezes)
Service Ping (all environments → data warehouse):
counts.count_ai_audit_events(28d + all) — PostgreSQLDatabaseMetric(timestamp_column :created_at, so the 28d frame prunes monthly partitions). Reports ~0 where ClickHouse is the store.counts.count_ai_audit_events_clickhouse(28d + all) — separateGenericMetricgated onglobally_enabled_for_analytics?, querying ClickHousecount()inline (noclickhousedata source exists in Service Ping;data_source: database).counts.count_total_ingest_ai_audit_events_batch_monthly(28d) —internal_eventscount of ingest batches.settings.ai_audit_events_streaming_enabled
Internal Event (all environments → Snowplow/Snowflake):
ingest_ai_audit_events_batch—value(batch size) andproperty(storage backend:clickhouse/postgresql), for cross-environment throughput visibility. Not fired for empty batches.
Why
ai_audit_events is routed to ClickHouse on GitLab.com and PostgreSQL on self-managed/Dedicated (ClickHouse is optional on self-managed). Both grow unbounded today, and rows are large (prompt content and metadata). Before committing to partitioning, sharding, truncation, or a retention policy we need to measure size/throughput and prove coverage in every environment. Service Ping (counts) + Internal Events (throughput) are the only signals that reach the warehouse from self-managed/Dedicated; Prometheus is the GitLab.com high-fidelity supplement. The Rails-side ingested/stored/buffered counters plus the producer-side captured counter (companion ai-gateway MR) form a captured ≈ ingested ≈ stored/buffered reconciliation that proves no loss (divergence is the loss signal, which is why ingested is counted pre-persistence and kept separate from stored/buffered).
Table size & growth observability
Row counts alone under-describe a table with large rows that grows exponentially. Size/growth is observed via existing infrastructure (no new metric needed here):
- PostgreSQL:
Gitlab::Database::PostgresTableSize/postgres_table_sizes(total/table/index/toast bytes,small/medium/large/over_limit, 25 GB alert) and thepg_total_relation_size_bytes{relname="ai_audit_events"}Prometheus series.ai_audit_eventsis already tracked there. - ClickHouse:
system.parts(sum(bytes_on_disk),sum(rows)per partition) plus the ClickHouse Prometheus endpoint. - Per-partition size/rows and average row size (
total_bytes / row_count) are the inputs to the partition/shard/TTL/truncation decisions — captured as the dashboard follow-up below.
Testing
- Ingest service, process-events worker, streaming worker,
write_buffer, and instrumentation metric specs all pass. - Metric and internal-event definition validation specs pass.
- Note: the
:click_house-taggeddump_write_buffer_workerexamples may fail locally with ClickHouseMEMORY_LIMIT_EXCEEDED(Code 241) — environmental, not a code issue; the gauge logic is also covered by the Redis-levelwrite_buffertests. They run in CI.
Reviewer notes
count_ai_audit_eventsis now a PostgreSQLDatabaseMetric(SQL generated on GitLab.com and executed later by the Data team), and the ClickHouse count is a separatecount_ai_audit_events_clickhouseGenericMetric. This keeps values consistent regardless of whether an instance toggles ClickHouse analytics. Open question for analytics-instrumentation: keep the CHGenericMetric, or drop it if the Data team counts ClickHouse directly?data_source: databaseis used because Service Ping has noclickhousesource.- The Service Ping
all-frame PostgreSQL count queriesai_audit_events(potentially large) —databasereview requested. gitlab_clickhouse_write_buffer_pendinglives in the sharedClickHouse::DumpWriteBufferWorker(feature_category: value_stream_management) — please loop in VSM/ClickHouse reviewers.- No migration. No changelog: internal observability, no user-facing change.
Out of scope (follow-ups)
- Retention/TTL + payload truncation (held pending the compliance decision on the retention window).
- Partitioning/sharding changes for unbounded growth.
- Promoting
buffered_totalto a true CH-persisted count (viaDumpWriteBufferWorkerinserted_rows). - Grafana dashboard for per-partition size/rows and average row size (PG + ClickHouse) and alert thresholds for
pg_fallback_total/write_buffer_pending(runbooks repo).
Database:
- ALL query: https://console.postgres.ai/gitlab/gitlab-production-main/sessions/52978/commands/154660
- Interval query: https://console.postgres.ai/gitlab/gitlab-production-main/sessions/52978/commands/154661