Add AI audit event storage adoption Service Ping metrics
What does this MR do and why?
Work item 607935 asks for Service Ping metrics covering AI audit event activity, including:
- Number of namespaces (groups/projects) with AI audit event storage enabled
This MR adds both sides of that count:
| Metric | Time frame | How |
|---|---|---|
settings.count_namespaces_with_ai_audit_events_storage_enabled |
none | batched count of namespace_settings rows with the flag on |
settings.count_projects_with_ai_audit_events_storage_enabled |
none | batched count of project_settings rows with the flag on |
settings.ai_audit_events_storage_enabled |
none | instance level application setting, read directly |
Event volume metrics cannot distinguish broad adoption from one busy namespace, so these count who has the storage toggle on, at group and project level separately. The instance level boolean mirrors the existing settings.ai_audit_events_streaming_enabled and lets analysts tell one admin enabling the whole instance apart from groups opting in, since the instance toggle cascades into every settings row.
Note on what a raw column count means here: ai_audit_events_storage_enabled is one of the cascading Duo settings. Changing it on a group enqueues Namespaces::CascadeDuoSettingsWorker, which writes the value into the namespace_settings and project_settings rows of all existing descendants, and new projects copy the group value at creation (ee/app/services/ee/projects/create_service.rb). So the counts reflect enablement materialized across the hierarchy, not just individual toggles.
Known gaps: subgroups created after the cascade keep a NULL column even though the cascading reader resolves them to enabled, and lock-only changes do not rewrite descendant rows. It only affects the namespaces metric accuracy, I suggest this is acceptable signal, because we count project and instance level settings as well.
References
Database review
Two post-deploy migrations add partial indexes so the batch counters can compute MIN/MAX and per-window counts on the filtered relation instead of scanning the full table. I created the indexes on the Database lab clone.
The batch counter computes its bounds directly on the filtered relation, then counts rows per 100k id window:
SELECT MIN("namespace_settings"."namespace_id") FROM "namespace_settings"
WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE
SELECT MAX("namespace_settings"."namespace_id") FROM "namespace_settings"
WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE
SELECT COUNT("namespace_settings"."namespace_id") FROM "namespace_settings"
WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE
AND "namespace_settings"."namespace_id" BETWEEN $start AND $finish
SELECT MIN("project_settings"."project_id") FROM "project_settings"
WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE
SELECT MAX("project_settings"."project_id") FROM "project_settings"
WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE
SELECT COUNT("project_settings"."project_id") FROM "project_settings"
WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE
AND "project_settings"."project_id" BETWEEN $start AND $finishPlans below are from Database Lab against a fresh clone of gitlab-production-main, with the indexes in place. All eight are index only scans on the new partial indexes. None touch the primary key index, and none seq scan.
| Query | Total time | Plan |
|---|---|---|
| MIN namespace_settings | 1.1 ms | 158414 |
| MAX namespace_settings | 1.1 ms | 158415 |
| MIN project_settings | 1.3 ms | 158416 |
| MAX project_settings | 1.2 ms | 158417 |
| namespaces window BETWEEN 9970 AND 109970 | 1.2 ms | 158418 |
| namespaces window BETWEEN 125000000 AND 125100000 | 1.2 ms | 158419 |
| projects window BETWEEN 278964 AND 378964 | 1.2 ms | 158420 |
| projects window BETWEEN 80000000 AND 80100000 | 1.7 ms | 158421 |
Adopters of this setting exist near both ends of each table's id range, so the filtered MIN/MAX nearly match the full id range of the table. That works out to about 1,350 windows for namespaces and about 950 for projects, well under the batch counter's 10,000 loop fallback cap. Empty windows are cheap, roughly 1 ms index probes each, so the walk is dominated by the batch counter's built-in 10 ms sleep between batches: about 15 seconds for the namespaces metric and about 10 seconds for the projects metric per Service Ping run.
Database Lab hides result rows, so the bounds below were bracketed with probe counts read from the plans in session 54763:
| Table | MIN | MAX | Batches at 100k batch size. Corresponds to count metric | Loop guard |
|---|---|---|---|---|
namespace_settings.namespace_id |
below 100 (89 rows under 100) | about 130M (5.5M rows at >= 125M, none at >= 140M) | ~1,350 | 7x under the 10,000 loop fallback |
project_settings.project_id |
below 100 (55 rows under 100) | about 90M (4.5M rows at >= 80M, none at >= 100M) | ~950 | 10x under |
How to set up and validate locally
require_relative 'spec/support/helpers/service_ping_helpers.rb'
ServicePingHelpers.get_current_usage_metric_value('settings.count_namespaces_with_ai_audit_events_storage_enabled')MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.