feat: add group dimension to AiUsageEvents and DuoWorkflows analytics

What does this MR do and why?

Adds group as a dimension and sort field to both analytics sources that expose it, so a query can bucket results per group:

dimensions: group
metrics: usersCount, creditsUsedSum
sort: usersCount desc

This is the DAP Impact v1 "Group comparison" panel, which today needs one request per group and is capped at 20 groups. It becomes a single query covering any number of groups.

Two compiler gaps blocked it, and both are fixed here:

  • No scalar integer parameter kind. ParameterConstraint had Enum, Range (f64) and List; #213 (closed) added ListItem::Int for integers inside a list, but a scalar Int argument had nowhere to live. A Range-backed default renders depth: 1.0, which the GraphQL Int argument rejects, so plain dimensions: group would have been broken while group(depth=2) worked. Adds ParameterConstraint::Int { min, max } and ParameterDefault::Int, mirroring ListItem::Int and reusing its parse_graphql_int grammar and validate_int_item bounds check rather than restating them.

  • A dimension could be parameterised or an object, never both. render_parameterized_body emitted name(args) and never consulted render_query_field, the only path that produced a selection set, so group(depth: 1) { … } was inexpressible. Adds a field_selection_set hook on SourceAnalyzer, defaulting to None, consulted on every path a field can take — parameterised, the empty-parameters fallback, and the plain Static arm. A source declares an object field's shape once.

depth mirrors the engine: absolute, counted from the hierarchy root rather than the query scope, default 1, bounds 1..99. Sorting needed no new code, because sort resolution already inherits the selected dimension's parameters — which matters, since the engine keys the row on the depth and an orderBy without parameters: { depth: N } fails with the specified identifier is not available: 'group'.

Why group is not a filter

The engines expose a groupId argument, but GLQL does not compile a filter to it. descendantsScope (group in (...), shipped in 19.3) already narrows to any set of groups and projects with the same subtree semantics, so a group filter would be a second spelling for the same thing. group = "path" also already means scope, so adding a filter under the same keyword would leave quoting to decide the mechanism.

group therefore keeps the meanings it has, plus one:

GLQL Compiles to Meaning
group = "gitlab-org" group(fullPath: "gitlab-org") query scope
group in ("a", "b") descendantsScope: {groupFullPaths: [...]} multi-group scope
dimensions: group group(depth: N) { … } breakdown (new)

Engine side

Both engines declare the dimension with the same framework class, so the GraphQL surface is identical across the two sources:

Example Usage

Build prerequisite, once:

cd glql_rb
bundle install
bundle exec rake compile

Group comparison panel (DuoWorkflows)

mode: analytics
query: type = duoworkflow
dimensions: group
metrics: usersCount, creditsUsedSum
sort: usersCount desc
Test via ruby extension
cd glql_rb # after the build prerequisite above

bundle exec ruby -Ilib - <<'RUBY'
require "gitlab_query_language"
require "json"
require "net/http"

query   = "type = duoworkflow"
context = { mode: "analytics", group: "gitlab-org", dimensions: "group",
            metrics: "usersCount, creditsUsedSum", sort: "usersCount desc" }

compiled = Glql.compile(query, context)
puts "== Generated GraphQL =="
puts compiled["output"]

uri  = URI("https://gitlab.com/api/graphql")
http = Net::HTTP.new(uri.host, uri.port); http.use_ssl = true
req  = Net::HTTP::Post.new(uri, "Content-Type" => "application/json")
req["PRIVATE-TOKEN"] = ENV.fetch("GITLAB_TOKEN")
req.body = { query: compiled["output"], variables: { limit: 3 } }.to_json
response = JSON.parse(http.request(req).body)

transformed = Glql.transform(response["data"], { mode: context[:mode], fields: compiled["fields"] })
puts "\n== Transformed rows =="
puts JSON.pretty_generate(transformed["data"])
RUBY

Generated GraphQL:

query GLQL($before: String, $after: String, $limit: Int) {
  group(fullPath: "gitlab-org") {
    analytics {
      duoWorkflows {
        aggregated(before: $before, after: $after, first: $limit, orderBy: [{direction: DESC, identifier: "usersCount"}]) {
          count
          pageInfo { startCursor endCursor hasNextPage hasPreviousPage }
          nodes {
            dimensions {
              group(depth: 1) { id fullPath webUrl fullName }
            }
            usersCount
            creditsUsed { creditsUsedSum: sum }
          }
        }
      }
    }
  }
}

Rows are not included: duoWorkflows is behind the dap_impact_v1 feature flag and needs a token with access to Duo analytics, so the response is not reproducible from a public request. The compile and transform stages are covered by tests, including a case asserting that a group: null row (a flow tracked directly in a project) survives flattening rather than being dropped.

Subgroup breakdown (AiUsageEvents)

mode: analytics
query: type = aiusageevent
dimensions: group(depth=2)
metrics: usersCount
sort: usersCount desc
Test via ruby extension

Same snippet as above, with:

query   = "type = aiusageevent"
context = { mode: "analytics", group: "gitlab-org", dimensions: "group(depth=2)",
            metrics: "usersCount", sort: "usersCount desc" }

Generated GraphQL:

query GLQL($before: String, $after: String, $limit: Int) {
  group(fullPath: "gitlab-org") {
    analytics {
      duoUsageEvents {
        aggregated(before: $before, after: $after, first: $limit, orderBy: [{direction: DESC, identifier: "usersCount"}]) {
          count
          pageInfo { startCursor endCursor hasNextPage hasPreviousPage }
          nodes {
            dimensions {
              group(depth: 2) { id fullPath webUrl fullName }
            }
            usersCount
          }
        }
      }
    }
  }
}

depth: 2 is absolute, so under a top-level group it yields one row per direct subgroup. Events tracked directly in a project at that depth resolve to group: null, which is engine behaviour and documented on the dimension; charts should drop those rows.

How to set up and validate locally

  1. cargo test — 1714 examples pass.
  2. cargo clippy --all-targets -- -D warnings and cargo fmt --check — clean.
  3. cargo run --bin generate-schema && npm run lint:prettier:fix — src/schema/schema.json regenerates with no drift.
  4. DUMP_GRAPHQL=1 cargo test && npm run test:graphql — every group document validates against the current schema dump.

Verification beyond the test suite

  • No regression in existing sources. A corpus of 16 queries covering all 7 analytics sources, every parameter kind (granularity enum, quantile float, thresholds list) and standard mode was compiled on main and on this branch. The generated GraphQL is byte-identical, so neither the new Int variant nor routing the Static arm through the hook perturbs existing rendering.
  • Engine conformance. Every dimension and metric GLQL publishes, for all 7 sources, was compiled and validated against the live GitLab schema. All valid.
  • Frontend, tables: the dimension resolves to a Group, which presentersByObjectType already renders with LinkPresenter (it needs webUrl and fullName, both in the selection), and the selection carries id so rows pair across periods for the trend column.
  • Frontend, charts: not covered. labelByObjectType in chart_data.js handles only UserCore and Project, so a Group renders an empty label — and because that string is the series identity downstream, distinct groups collapse into one entry. A follow-up in gitlab-org/gitlab must register a Group formatter before this dimension is charted. Raised by @jiaan below.
  • Transformer untouched. src/transformer/context.rs is byte-identical to main.

Follow-ups

Neither blocks this MR, both are in gitlab-org/gitlab:

  • Chart formatter. Register Group in labelByObjectType, e.g. (value) => value.fullPath ?? value.fullName. Required before the dimension is used in a chart, on either source.
  • User docs. doc/user/glql/data_sources/duo_workflows.md and ai_usage_events.md list every dimension and neither has a group row or the depth parameter. duo_workflows.md is also missing userTier, so one follow-up could cover both gaps.

Issue references

Closes #169 (closed), which covers both AiUsageEvents and DuoWorkflows.

Follow-up #228 (closed) moves the other object fields (project, user, pipeline, …) onto the same hook, after this and the #187 (closed) train land.

Edited by Chandra Saripaka

Merge request reports

Loading
Loading