Add event name filtering to the AI usage events total count

What does this MR do and why?

Adds an optional event parameter to the count metric of the AI usage events aggregation engine, so one GraphQL request can return an unfiltered total alongside one or more event-filtered totals:

totalCount
chatCount: totalCount(event: "request_duo_chat_response")
chatOrShownCount: totalCount(event: ["request_duo_chat_response", "code_suggestion_shown_in_ide"])

Previously a dashboard showing a headline total next to a per-event breakdown had to issue one extra query per event and stitch the results together client-side.

Passing several event names ORs them. Passing an unrecognized event name returns an error instead of a silently wrong count.

This follows the pattern two other engines already use: Pipelines (parameters source and status) and DuoWorkflows (parameter status). The generic aggregation framework needed no changes.

References

Related issue: https://gitlab.com/gitlab-org/gitlab/-/issues/629550

Notes for reviewers

  • The filter and the new parameter handle unknown event names differently. The pre-existing event filter silently ignores unrecognized event names. The new event metric parameter added here rejects them with an error, which is what the issue's acceptance criteria asked for. Fixing the filter is out of scope for this MR.

  • This does not prune ClickHouse granules. That's inherent to the feature.

    Why

    A metric parameter isn't a WHERE predicate; it compiles to a conditional aggregate, countIf(...), evaluated over the rows already in scope. The table is ordered by (traversal_path, event, timestamp, user_id), so the existing event filter prunes hard on the second sort-key column, while totalCount(event: X) scans the whole traversal_path range. That's the cost of getting the total and the per-event breakdown in a single pass, which is the point of this feature. A caller who only wants the filtered number should use the filter instead.

    GraphQL query complexity (250 for authenticated requests) is the only cap on how many aliased counts one request can ask for. There are 44 registered events; roughly 44 aliased totalCount fields score about 90 in complexity, so a client can legitimately request every event name at once and get back about 44 countIf columns.

  • Two pre-existing framework gaps are not fixed here. Both are tracked in #630494

    Why

    First, metric parameters can't carry an experiment: marker, because declare_parameter_arguments forwards only type, array, and description to the GraphQL argument. So this argument lands as stable public API as soon as this merges, the same way the equivalent Pipelines and DuoWorkflows arguments did.

    Second, the generated argument has no explicit array-length validation of its own. It relies on the automatic 1000-item cap that Types::BaseArgument applies to every array argument. Both gaps affect all three engines that use parameters, not just this one.

How to set up and validate locally

  1. Enable ClickHouse for analytics in your GDK and seed some ai_usage_events rows.

  2. In GraphiQL, run a query against a group you can read:

    query {
      group(fullPath: "your-group") {
        analytics {
          duoUsageEvents {
            aggregated {
              nodes {
                totalCount
                chatCount: totalCount(event: "request_duo_chat_response")
              }
            }
          }
        }
      }
    }
  3. Confirm chatCount is the subset of totalCount matching that event, and that passing an unknown name such as totalCount(event: "nonexistent") returns an Invalid value(s) for parameter error.

Testing

  • Model-level specs cover: a single event, several events ORed, total and filtered counts requested together, a zero result, the parameter combined with a contradicting filter, grouping by the event dimension while filtering, and an invalid event name.
  • A GraphQL request spec covers the aliasing case and the invalid-name error, and runs for group, project, and organization scopes.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Chandra Saripaka

Merge request reports

Loading
Loading