Add exact_not_match filters to all aggregation engines

What does this MR do and why?

Every list filter across the eight analytics aggregation engines could only include values, never exclude them. That blocked common asks: excluding bot accounts from usage metrics, excluding the default branch from pipeline and deployment breakdowns, excluding skipped pipelines from duration numbers.

This adds an exact_not_match filter definition alongside the existing exact_match for 26 of the existing list filters, so each column also exposes a <name>_not filter, surfaced in GraphQL as <name>Not. Each exact_not_match is declared explicitly on its own line rather than derived from a shared DSL helper, per reviewer feedback and the issue's acceptance criteria. Formatters and expressions are shared with the positive filter, so no existing argument changes behavior. created_by_duo stays inclusion-only (see the table below).

Closes https://gitlab.com/gitlab-org/gitlab/-/issues/629296.

Note

The framework MR !256159 (merged) has merged, so this now targets master and every commit in the diff belongs to this MR. The newest commits are the response to the first review round: Declare exclusion filters explicitly instead of via a DSL helper, Forward the experiment option to generated GraphQL arguments, State the NOT IN gotchas on ExactNotMatchFilter and drop the spec copies, Pin the remaining NULL caveats and narrow the coverage claim, and the regenerated GraphQL artifacts.

Engine Filters paired
Pipelines status, source, ref, user_id
Deployments status, environment_id, ref, user_id
Merge Requests target_branch, state_id, author_id
Duo Workflows workflow_definition, status, project_id, user_id
AI Usage Events user_id, event, feature, flow_type
Agent Platform Sessions user_id, flow_type, project_id
Code Suggestions user_id, language, ide_name
Contributions author_id

created_by_duo on Merge Requests is deliberately left inclusion-only: it is a non-nullable boolean, so createdByDuo: false already expresses the exclusion. It is recorded in the inclusion_only allowlist in ee/spec/models/analytics/aggregation_engines/exclusion_filter_coverage_spec.rb.

Not marked experiment, and why

These arguments were going to ship as experiment, to give a rollback path. That is reverted: marking them made every one of them blow the stack at query time.

Gitlab::Graphql::Deprecations turns an experiment: marker into a deprecation_reason. On these resolver-generated arguments that causes Gitlab::Graphql::Tracers::InstrumentationTracer#execute_multiplex to recurse — its ensure block calls export_query_info, then error_type, then query.result, and when execution raises before the result is memoized, query.result re-enters execution and runs the tracer again. Reproduced with GitlabSchema.execute on all 26 arguments across all 8 engines; the positive filters, which carry no marker, are unaffected. It is the marker and nothing else: removing the single adapter line that forwards it fixes every argument, and restoring the line breaks them again.

This looks like a framework bug rather than something to solve in this MR, so the arguments ship unmarked, which is the behavior that was already green in CI. Tracked in #630513.

max_size is uniform: every exclusion filter uses MAX_EXCLUSION_VALUES (100), a named constant on the ClickHouse Engine, and each description states the limit. An earlier revision capped only the sort-key columns on the theory that it helped pruning, which the measurements below disprove — list size does not change how much an exclusion reads. The positive filters stay uncapped, since capping them would newly reject queries that work today.

Why: measured locally against ClickHouse with 2.5M rows per table, one part, 307 granules, index_granularity 8192, real sort keys, scoped to one group prefix (synthetic local data, not production). EXPLAIN indexes = 1 granules selected:

  • ai_usage_events.event (sort position 2, 40 distinct values): IN selects 24/306 granules, NOT IN selects 286/306.
  • agent_platform_sessions.user_id (position 2, 5k distinct values): IN selects 1/306, NOT IN selects 306/306.
  • ai_code_suggestions.user_id (position 2, 5k distinct values): IN selects 1/306, NOT IN selects 306/306.

Read volume for user_id NOT IN was 2,500,000 rows / 30.99 MiB, versus 8,192 rows / 71.81 KiB for IN. A tighter max_size does not fix this: 3, 100, and 1000 excluded values all still selected 306/306 granules.

A max_execution_time guard is not in this MR. ClickHouse::Client.select takes no per-query settings, and per-database variables in config/click_house.yml would apply to every ClickHouse query in the app, so that is a framework-wide change that belongs in its own MR. Tracked in #630512.

Things worth a reviewer's attention

Two deliberate decisions, so they are declared rather than discovered:

  1. Two arguments drop most of their data, and that is intended. ai_usage_events.flow_type and .feature are NULL for non-Duo-Agent-Platform events and for unregistered features respectively, and NOT IN drops NULL rows. So flowTypeNot returns only Duo Agent Platform events. That matches the motivating use case in #629295 (closed) ("exclude a noisy flow type from Duo Agent Platform stats"), so they are paired rather than left inclusion-only. Engine owners should overrule this if consumers would read it differently. This is now documented as a class comment on ExactNotMatchFilter, and tracked for the API docs in #629847.

  2. An unrecognized exclusion value returns every row, and a spec pins that. Enum formatters discard values they cannot resolve, leaving an empty NOT IN, which Arel renders as 1=1. statusNot: ["typo"] therefore returns everything, where status: ["typo"] returns nothing. This predates the MR and already applies to the positive filters, which have their own specs asserting it. The green spec here is not an endorsement — it documents it. This is also now documented as a class comment on ExactNotMatchFilter. Fixing it properly means request-level validation, tracked in #629846.

Test coverage

!256159 (merged) covers a single exclusion filter's mechanics, including NULL handling at framework level. This adds the per-engine coverage and the interactions:

  • One negative example per identifier — 26 across the eight engine specs — plus one GraphQL request example per engine.
  • NULL handling for every nullable column an exclusion description mentions: pipelines.user_id, pipelines.ref, pipelines.source, deployments.user_id, merge_requests.author_id, duo_workflows.project_id, and ai_usage_events.flow_type / feature. Each asserts the NULL row is dropped, and each fails if it is kept. The feature case needs a stub: the ELSE NULL branch is unreachable with real data, so that example narrows Gitlab::Tracking::AiTracking.registered_features.
  • Formatter-emptied lists, for pipeline sources and merge request states.
  • Interactions: grouping by a column that is also excluded, two exclusions on different columns, and an exclusion combined with a range filter.
  • A drift guard, ee/spec/models/analytics/aggregation_engines/exclusion_filter_coverage_spec.rb. It asserts every exact_match filter has an exact_not_match twin on the same column, that every exclusion identifier ends in _not (aside from the inclusion_only allowlist, which currently holds created_by_duo), and that the engine list matches the files on disk. A ninth engine, or a new exact_match filter on an existing one, now fails the suite unless it is paired or recorded in that allowlist. The guard is scoped to exact_match: descendants and metric_exact_match have no NOT form to pair with, which the spec header and the developer docs both state.

The NULL and empty-list rules for exact_not_match are now documented once, as a class comment on ExactNotMatchFilter, instead of being repeated across specs; nine duplicated spec comments restating them were removed.

How to set up and validate locally

Requires ClickHouse in GDK.

bundle exec rspec spec/lib/gitlab/database/aggregation ee/spec/models/analytics/aggregation_engines

470 framework specs and 368 engine specs pass locally; RuboCop is clean.

The GraphQL request specs under ee/spec/requests/api/graphql/analytics still cannot be run locally: they exercise the full controller stack, which reaches universal_path_to_stylesheet and fails with LoadError: cannot load such file -- sass. Untouched examples in the same files fail identically, so this is a pre-existing local limitation, not something these changes introduce. CI runs them.

Argument wiring was instead verified with GitlabSchema.validate for all 26 arguments, which also confirmed createdByDuoNot is rejected and createdByDuo is still valid.

Then confirm in GraphiQL that an exclusion is the complement of its inclusion:

query {
  group(fullPath: "gitlab-org") {
    analytics {
      pipelines(statusNot: ["skipped", "canceled"]) {
        aggregated { nodes { dimensions { ref } totalCount } }
      }
    }
  }
}

Generated files

doc/api/graphql/reference/_index.md and public/-/graphql/introspection_result.json are regenerated and isolated in their own commits. introspection_result_no_deprecated.json is deliberately unchanged: analytics is an experiment, which is modeled as a deprecation, so the whole Analytics type is absent from that file.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Chandra Saripaka

Merge request reports

Loading
Loading