Add negative filters (!= / not in) on analytics list filters

What does this MR do and why?

Analytics list filters can only include values, never exclude them. The DAP Impact v1 Spend panel "Credits over time excluding Chat" wants every Duo flow type except chat and agentic_chat/v1, so it ships an explicit include list (gitlab!255865 (merged)) that goes stale the moment a new flow type appears.

The aggregation engines now expose an exact_not_match twin for every list filter as a <name>Not GraphQL argument. This adds the GLQL side: a not in operator, and negation on all 17 analytics filter fields whose engine column got a twin.

flowType not in ("chat", "agentic_chat/v1")
  -> duoWorkflows(workflowDefinitionNot: ["chat", "agentic_chat/v1"])

Closes #221 (closed).

The spelling, settled once

#221 (closed) asks for the operator set to be decided for the whole language before building the mechanism, and #69 asks for not in on Status. This MR settles it: not in is the list-valued negation, and != stays scalar-only.

That mirrors the existing pair exactly — a scalar spelling and a list spelling on each side:

scalar list
include flowType = "chat" flowType in ("chat", "web")
exclude flowType != "chat" flowType not in ("chat", "web")

The alternative — teaching != to take a list — was rejected because = rejects a list, so it would have introduced an asymmetry rather than removing one.

Both negated spellings resolve to the same <name>Not argument, because transform_expression collapses not in onto != for has-one lists. That mirrors the rule that already collapses in onto =, which is why analytics in compiles flat rather than into or: { … }.

Fields covered

17 GLQL fields across 7 sources, 13 distinct argument names (userIdNot recurs on four sources, statusNot on two):

Source GLQL field → engine argument
Duo Workflows flowType → workflowDefinitionNot, status → statusNot, user → userIdNot
Agent Platform Sessions flowType → flowTypeNot, user → userIdNot
AI Usage Events feature → featureNot, event → eventNot, user → userIdNot
Code Suggestions language → languageNot, ideName → ideNameNot, user → userIdNot
Contributions user → authorIdNot
Merge Requests state → stateIdNot, targetBranch → targetBranchNot
Pipelines status → statusNot, ref → refNot, pipelineSource → sourceNot

createdByDuo on Merge Requests stays inclusion-only, matching the engine: it is a non-nullable boolean, so createdByDuo: false already expresses the exclusion, and transform_expression keeps flipping != true to = false. A test pins that it does not become createdByDuoNot.

No GLQL work for three engine twins: duo_workflows.project_id and agent_platform_sessions.project_id are query scope in GLQL rather than filters, and GLQL's AI Usage Events source exposes no flowType filter.

How negation reaches a flat argument

Codegen picks a filter's map from its operator: = and the range operators go to the flat and map, in to or: { … }, != to not: { … }. Analytics engine fields have no not: argument. The engine carries the negation in the argument name instead.

Whether that happens is a property of the argument, not of the source: on one engine, one column can have an exact_not_match twin while its neighbour does not. So the fact travels with the key. SourceAnalyzer::filter_key now returns the FilterKey and graphql_filter_key becomes a default method that resolves it, so codegen can ask FilterKey::absorbs_negation() when choosing the map.

That refactor is isolated in its own commit and changes no output: absorbs_negation() is false for all four pre-existing variants, and the generated schema is byte-for-byte unchanged at that commit.

Three guards keep it honest:

  • test_analytics_negation_is_keyed_into_the_argument_name — every analytics field admitting != or not in must resolve to a Not-suffixed key, unless it is boolean-like, where the negation is flipped away before codegen. Without it, a field declared as a bare StringLike (which admits != by default) would be emitted flat under its positive name, silently inverting the filter. Verified it fails when flowType is reverted to a non-negatable key.
  • test_analytics_negatable_keys_belong_to_negatable_fields — the converse. A Negatable key on a field whose type admits neither negating operator is dead configuration that reads as supported while every query using it is rejected at the operator gate.
  • analytics::generate refuses to emit a query when the or or not map is non-empty, naming the offending key. Both tests inspect declarations; this is the output-level backstop for a shape no test happens to compile.

Verified end to end against gitlab.com

gitlab!256398 (merged) deployed on 2026-09-23, so the full compile → execute → transform cycle runs against production. Grouping by status on gitlab-org/gitlab partitions the pipelines, which makes the exclusion checkable as a set complement rather than just "the argument was accepted":

Query Statuses returned
no status filter canceled, canceling, created, failed, manual, running, skipped, success
status = success success
status != success canceled, canceling, created, failed, manual, running, skipped

So the negated result is exactly the unfiltered set minus success, and the two are disjoint. All three stages pass: the query compiles to pipelines(statusNot: ["success"]), the API returns rows, and Glql.transform flattens them.

One caveat for anyone repeating this with a non-partitioning dimension: aggregated.count counts dimension groups, not rows, so summing the positive and negated counts over a dimension like ref exceeds the unfiltered total. A ref can have both successful and failed pipelines and therefore appears in both sets. That is expected, not a filter bug.

duoWorkflows additionally sits behind the dap_impact_v1 feature flag, so the Spend panel needs that enabled before its queries return rows.

Example Usage

Build prerequisite, once:

cd glql_rb
bundle install
bundle exec rake compile

Excluding two flow types — the Spend panel case

mode: analytics
query: type = DuoWorkflow and group = "gitlab-org" and created > -6m and flowType not in ("chat", "agentic_chat/v1")
dimensions: created(monthly), flowType
metrics: creditsUsedSum
Test via ruby extension
cd glql_rb # after the build prerequisite above

bundle exec ruby -Ilib - <<'RUBY'
require "gitlab_query_language"
require "json"
require "net/http"

query   = 'type = DuoWorkflow and group = "gitlab-org" and created > -6m and flowType not in ("chat", "agentic_chat/v1")'
context = { mode: "analytics", group: "gitlab-org", dimensions: "created(monthly), flowType", metrics: "creditsUsedSum" }

# 1. Compile GLQL -> GraphQL (also returns the typed `fields` the transform needs).
compiled = Glql.compile(query, context)
puts "== Generated GraphQL =="
puts compiled["output"]

# 2. Execute against GitLab. Analytics is private data, so a token is required:
uri  = URI("https://gitlab.com/api/graphql")
http = Net::HTTP.new(uri.host, uri.port); http.use_ssl = true
req  = Net::HTTP::Post.new(uri, "Content-Type" => "application/json")
req["PRIVATE-TOKEN"] = ENV.fetch("GITLAB_TOKEN")
req.body = { query: compiled["output"], variables: { limit: 3 } }.to_json
response = JSON.parse(http.request(req).body)

# 3. Transform the response back into flat rows.
transformed = Glql.transform(response["data"], { mode: context[:mode], fields: compiled["fields"] })
puts "\n== Transformed rows =="
puts JSON.pretty_generate(transformed["data"])
RUBY

Generated GraphQL:

query GLQL($before: String, $after: String, $limit: Int) {
  group(fullPath: "gitlab-org") {
    analytics {
      duoWorkflows(createdAtFrom: "2026-03-23 23:59", workflowDefinitionNot: ["chat", "agentic_chat/v1"]) {
        aggregated(before: $before, after: $after, first: $limit) {
          count
          pageInfo { startCursor endCursor hasNextPage hasPreviousPage }
          nodes {
            dimensions {
              createdAt(granularity: "monthly")
              flowType: workflowDefinition
            }
            creditsUsed { creditsUsedSum: sum }
          }
        }
      }
    }
  }
}

Two exclusions on one source, and the scalar spelling

mode: analytics
query: type = Pipeline and project = "gitlab-org/gitlab" and status != success and ref not in ("main", "master")
dimensions: ref
metrics: totalCount
Test via ruby extension

Same snippet as above with:

query   = 'type = Pipeline and project = "gitlab-org/gitlab" and status != success and ref not in ("main", "master")'
context = { mode: "analytics", project: "gitlab-org/gitlab", dimensions: "ref", metrics: "totalCount" }

Generated GraphQL (engine field arguments only):

pipelines(statusNot: ["success"], refNot: ["main", "master"]) {

Both spellings land on their own argument, and status keeps its lowercasing.

Inclusion and exclusion side by side

mode: analytics
query: type = DuoWorkflow and group = "gitlab-org" and status in (finished, failed) and user != 12345
dimensions: status
metrics: totalCount
Test via ruby extension

Generated GraphQL (engine field arguments only):

duoWorkflows(status: ["finished", "failed"], userIdNot: ["gid://gitlab/User/12345"]) {

They are distinct arguments, so both survive. user keeps its User-GID mapping on the negated side.

Numeric-ID exclusion, GID mapping preserved

mode: analytics
query: type = Contribution and group = "gitlab-org" and user not in (1, 2)
dimensions: created
metrics: totalCount
Test via ruby extension

Generated GraphQL (engine field arguments only):

contributions(authorIdNot: ["gid://gitlab/User/1", "gid://gitlab/User/2"]) {

The engine names this column author_id, so the twin is authorIdNot, not userIdNot.

An empty exclusion list is rejected

mode: analytics
query: type = DuoWorkflow and group = "gitlab-org" and flowType not in ()
dimensions: flowType
metrics: totalCount
Test via ruby extension
Error: `flowType not in (...)` needs at least one value to exclude. Excluding
nothing matches every row, so an empty list is almost certainly not what was
meant.

NOT IN () renders as 1=1 on the engine, so an empty exclusion would silently return every row — the opposite of what someone writing an exclusion wants, and invisible on a dashboard. Rejecting is better than compiling it, and better than silently dropping the expression, which would make a typo indistinguishable from no filter at all.

The positive counterpart flowType in () is deliberately left alone: it compiles to IN (), which matches nothing. Same shape, harmless direction, and already shipped.

Things worth a reviewer's attention

Two exclusion caveats are recorded next to the field types they affect, because the argument name cannot convey them:

  • feature on AI Usage Events is derived with a CASE … ELSE NULL END, and NOT IN drops NULL rows, so feature != x returns only events whose feature resolved — not "everything except x". Deliberate on the engine side; API wording tracked in gitlab#629847.
  • ref and pipelineSource on Pipelines are nullable for the same reason.

An unrecognized exclusion value can invert the filter, and GLQL only guards part of it. The engine's formatters drop values they cannot resolve, leaving an empty NOT IN that matches every row — so source != "psuh" returns every pipeline, where source = "psuh" correctly returns none. GLQL's closed enums already prevent this for status on Pipelines and Duo Workflows and for state on Merge Requests, each with a test.

pipelineSource is the one that could be closed — GitLab's schema has an authoritative enum CiPipelineSources with 20 values. It is left open here deliberately: adding the enum would newly reject source = "<new value>" whenever GitLab adds a source and GLQL has not caught up, which is the same staleness this MR exists to remove. Worth a follow-up decision rather than a silent change to an existing filter. Tracked backend-side in gitlab#629846.

A pre-existing silent drop, now reachable. combine_values handled Quoted, Reference and List only, so numbers and bare tokens fell to an error that combine_expressions swallows — user != 1 and user != 2 compiled to userIdNot: [".../User/1"], returning rows the author had excluded. The unwrap_or predates this branch, but those queries used to be rejected at the operator gate, so admitting negation turned a rejection into a wrong answer. Fixed, and the fix also unions the inclusion side (user = 1 and user = 2), which is a behaviour change outside the negation feature proper — consistent with how quoted strings already behaved.

Open question for reviewers: are ai_code_suggestions.language and ide_name nullable? gitlab!256398 (merged)'s NULL-handling specs cover pipelines.ref/source, ai_usage_events.feature/flow_type, merge_requests.author_id and duo_workflows.project_id, but not these two, which suggests they are not. If they are, language != "go" also drops rows with no language and the same caveat should be recorded on that field type.

List size is already capped. MAX_ARRAY_SIZE (100) in analyzer/value.rs applies to every list value, which matches the MAX_EXCLUSION_VALUES (100) cap the engine applies to exclusion filters. A test pins that 101 exclusion values are refused.

Test coverage

1850 tests pass; clippy --all-targets -- -D warnings and cargo fmt --check are clean; src/schema/schema.json regenerates with no drift; wasm-pack build --target web succeeds on the current head.

The 17 fields collapse to five equivalence classes — the only things that distinguish a field's path are its FilterKey variant, its field type and its value shaper — so the matrix is covered on the axes rather than by repeating fields.

  • negation_matrix_tests.rs sweeps the registry: it asks every analytics analyzer which filters admit a negating operator, derives a sample value from the field type, and drives both spellings through a real compile, asserting the Not-suffixed argument and no or:/not: wrapper. It names no field and no argument, so a new negatable field is swept the day it is declared. The swept count is asserted, because an empty sweep would satisfy every other assertion while proving nothing.
  • Per-source acceptance tests for all 17 fields in both spellings, each asserting the output carries no not: and no or: block.
  • Value axis: not in () rejected; 101 values refused; two lists merge and de-duplicate; the same value twice collapses; != (), = (), not in (ANY), != ANY, != NONE and != null all rejected, with the error kind asserted since they fail at different gates.
  • Polarity axis: in + not in, = + !=, = + not in, in + !=. Repeated exclusions union across quoted strings, numbers and tokens, in both scalar/list orders.
  • Naming axis: an odd-cased field resolves to the canonical argument. NegatableFieldName interpolates expr.field, so this only holds because transform_expression dealiases first.
  • Interactions: a scope list beside an exclusion (project in (...) must not reach the or map, or the wrapper guard would reject a valid query), an exclusion beside a metric-range filter that relocates onto aggregated, and an exclusion beside a sort.
  • Eight existing "rejects !=" tests converted from rejection to acceptance. The standard-mode rejection tests are deliberately untouched and still green, which is the regression signal that scalar_or_list_field_type was left alone for standard mode.
  • schema_conformance.rs sweeps both new spellings against every field of every source, in both directions — advertised operators must compile, unadvertised ones must be rejected.

Schema validation

DUMP_GRAPHQL=1 cargo test && npm run test:graphql over 1711 dumped queries: all documents are valid against the published master schema.

Dependencies

Both engine-side MRs have merged and deployed:

  • #69 — in, not in and != on Status. This MR settles the not in spelling; extending it to standard-mode sources is a follow-up.
  • #173 — every operator on type compiles silently, the failure mode the registry-wide invariant test is designed to avoid elsewhere.
  • #188 — tracking issue.
Edited by Chandra Saripaka

Merge request reports

Loading
Loading