Add AI audit event storage adoption Service Ping metrics

What does this MR do and why?

Work item 607935 asks for Service Ping metrics covering AI audit event activity, including:

  • Number of namespaces (groups/projects) with AI audit event storage enabled

This MR adds both sides of that count:

Metric Time frame How
settings.count_namespaces_with_ai_audit_events_storage_enabled none batched count of namespace_settings rows with the flag on
settings.count_projects_with_ai_audit_events_storage_enabled none batched count of project_settings rows with the flag on
settings.ai_audit_events_storage_enabled none instance level application setting, read directly

Event volume metrics cannot distinguish broad adoption from one busy namespace, so these count who has the storage toggle on, at group and project level separately. The instance level boolean mirrors the existing settings.ai_audit_events_streaming_enabled and lets analysts tell one admin enabling the whole instance apart from groups opting in, since the instance toggle cascades into every settings row.

Note on what a raw column count means here: ai_audit_events_storage_enabled is one of the cascading Duo settings. Changing it on a group enqueues Namespaces::CascadeDuoSettingsWorker, which writes the value into the namespace_settings and project_settings rows of all existing descendants, and new projects copy the group value at creation (ee/app/services/ee/projects/create_service.rb). So the counts reflect enablement materialized across the hierarchy, not just individual toggles.

Known gaps: subgroups created after the cascade keep a NULL column even though the cascading reader resolves them to enabled, and lock-only changes do not rewrite descendant rows. It only affects the namespaces metric accuracy, I suggest this is acceptable signal, because we count project and instance level settings as well.

References

Database review

Two post-deploy migrations add partial indexes so the batch counters can compute MIN/MAX and per-window counts on the filtered relation instead of scanning the full table. I created the indexes on the Database lab clone.

The batch counter computes its bounds directly on the filtered relation, then counts rows per 100k id window:

SELECT MIN("namespace_settings"."namespace_id") FROM "namespace_settings"
  WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE

SELECT MAX("namespace_settings"."namespace_id") FROM "namespace_settings" 
  WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE

SELECT COUNT("namespace_settings"."namespace_id") FROM "namespace_settings"
  WHERE "namespace_settings"."ai_audit_events_storage_enabled" = TRUE
  AND "namespace_settings"."namespace_id" BETWEEN $start AND $finish

SELECT MIN("project_settings"."project_id") FROM "project_settings"
  WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE

SELECT MAX("project_settings"."project_id") FROM "project_settings"
  WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE

SELECT COUNT("project_settings"."project_id") FROM "project_settings" 
  WHERE "project_settings"."ai_audit_events_storage_enabled" = TRUE 
  AND "project_settings"."project_id" BETWEEN $start AND $finish

Plans below are from Database Lab against a fresh clone of gitlab-production-main, with the indexes in place. All eight are index only scans on the new partial indexes. None touch the primary key index, and none seq scan.

Query Total time Plan
MIN namespace_settings 1.1 ms 158414
MAX namespace_settings 1.1 ms 158415
MIN project_settings 1.3 ms 158416
MAX project_settings 1.2 ms 158417
namespaces window BETWEEN 9970 AND 109970 1.2 ms 158418
namespaces window BETWEEN 125000000 AND 125100000 1.2 ms 158419
projects window BETWEEN 278964 AND 378964 1.2 ms 158420
projects window BETWEEN 80000000 AND 80100000 1.7 ms 158421

Adopters of this setting exist near both ends of each table's id range, so the filtered MIN/MAX nearly match the full id range of the table. That works out to about 1,350 windows for namespaces and about 950 for projects, well under the batch counter's 10,000 loop fallback cap. Empty windows are cheap, roughly 1 ms index probes each, so the walk is dominated by the batch counter's built-in 10 ms sleep between batches: about 15 seconds for the namespaces metric and about 10 seconds for the projects metric per Service Ping run.

Database Lab hides result rows, so the bounds below were bracketed with probe counts read from the plans in session 54763:

Table MIN MAX Batches at 100k batch size. Corresponds to count metric Loop guard
namespace_settings.namespace_id below 100 (89 rows under 100) about 130M (5.5M rows at >= 125M, none at >= 140M) ~1,350 7x under the 10,000 loop fallback
project_settings.project_id below 100 (55 rows under 100) about 90M (4.5M rows at >= 80M, none at >= 100M) ~950 10x under

How to set up and validate locally

require_relative 'spec/support/helpers/service_ping_helpers.rb'
ServicePingHelpers.get_current_usage_metric_value('settings.count_namespaces_with_ai_audit_events_storage_enabled')

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Vasyl Pedak

Merge request reports

Loading
Loading