Draft: SPIKE: Add group dimension and filter to AiUsageEvents aggregation engine
Depends on !250611 (closed) (branch
605518-traversal-path-dimension), which introduces thetraversal_pathdimension this MR builds on. This MR is rebased onto and targeted at that branch. Review and merge !250611 (closed) first.
What does this MR do and why?
The "Duo adoption by group" dashboard chart needs distinct AI users broken down per group. The Analytics::AggregationEngines::AiUsageEvents ClickHouse aggregation engine already computes the distinct users metric and the time bucketing, but it had no way to group results by namespace hierarchy. Its namespace_path column (a traversal path such as 9/12/34/) was only used for scoping, not for grouping.
This MR adds that grouping capability, in five commits:
- A new reusable framework piece in
lib/gitlab/database/aggregation/click_house/:HasAnyFilter(DSL keywordhas_any), which filters rows where an array-valued expression contains at least one of the given values, using ClickHouse'shasAny(). Thetraversal_pathdimension used in commit 3 below is introduced separately in !250611 (closed). - GraphQL support for parameterized association dimensions. Dimension parameters are now declared as field arguments and included in the result-row key lookup.
declare_association_fieldmoved fromEngineResponseDimensionsTypeintoBaseResponseType, next todeclare_parameterized_field. This also fixed a RuboCopMetrics/AbcSizeoffense on thebuildmethod. - Wiring in the
AiUsageEventsengine:- A
group_iddimension usingtraversal_pathonnamespace_path, with an association that resolves toGroup(GraphQL fieldgroup,Types::GroupType, authorized per object withread_group). Ids that are not groups, such as project namespaces, user namespaces, or the0unattributed bucket, resolve to a nullgroup. - A
groupIdfilter that accepts group Global IDs and matches events attributed to the group itself or any of its descendants, by checking path segment membership.
- A
- Developer docs for the
has_anyfilter and a note about parameterized association dimensions indoc/development/aggregation_engines.md, plus regenerated GraphQL reference docs. Thetraversal_pathdimension docs moved to !250611 (closed). - Depth is now resolved against the group being queried, and the
groupIdfilter checks the model behind each Global ID. Both are described below.
The depth argument is optional and relative
The group dimension resolves its depth against the scope the request runs on, so the common case takes no argument at all. Under a group, group returns its direct children; group(depth: 2) drills one level further. Clients no longer need to know where their group sits in the hierarchy to build a query, and the same query means the same thing on every dashboard. The mechanics live in !250611 (closed).
Example query:
query {
group(fullPath: "flightjs") {
analytics {
duoUsageEvents {
aggregated {
nodes {
dimensions {
group { id fullPath }
timestamp(granularity: "monthly")
}
usersCount
}
}
}
}
}
}Global ID type checking on the filter
Namespace ids are shared between Group, ProjectNamespace and UserNamespace, so the groupId formatter now checks the model name on each parsed Global ID. Without it, a Global ID for an unrelated model whose numeric id happens to collide with a group id would silently act as a group filter. Values that are not group Global IDs are dropped, and a filter left with no values matches nothing.
Generated SQL
For ClickHouse reviewers, this is the shape of SQL the new pieces generate. The dimension example is a request for the direct children of a top-level group:
-- Dimension inner projection
toUInt64OrNull(arrayElement(splitByChar('/', namespace_path), 2)) AS aeq_group_id
-- Filter
WHERE hasAny(splitByChar('/', namespace_path), array('123'))namespace_path is the first column of the table's ORDER BY key, and scoping still applies the existing startsWith(namespace_path, ...) prefix condition, so queries stay bounded to the requested groups and projects. Note that the hasAny() filter itself cannot use the primary key index, and relies on that scoping prefix to bound the read.
References
Implements https://gitlab.com/gitlab-org/gitlab/-/issues/605518
The has_any filter is built at the framework level on purpose, so it can be reused by other engines. The traversal_path dimension is introduced separately in !250611 (closed). Planned follow-ups:
- https://gitlab.com/gitlab-org/gitlab/-/issues/605525 (MergeRequests engine)
- #608747 (Deployments engine)
How to set up and validate locally
- Enable ClickHouse in GDK, see https://gitlab.com/gitlab-org/gitlab-development-kit/-/blob/main/doc/howto/clickhouse.md, and make sure
Gitlab::ClickHouse.globally_enabled_for_analytics?is true. You can set this withGitlab::CurrentSettings.update!(use_clickhouse_for_analytics: true). - You need an EE license with the
ai_analyticsfeature. - Seed some
Ai::UsageEventrecords for projects in a group hierarchy, for example from a rails console. Then either letClickHouse::DumpWriteBufferWorkerand the sync workers run, or insert rows into theai_usage_eventsClickHouse table directly. - Run the example GraphQL query above in GraphiQL (
/-/graphql-explorer) against a group you can read, and check that results are bucketed by that group's direct children and that thegroupIdfilter argument works. Running the same query against a subgroup should bucket by the subgroup's own children. - Or run the specs:
bundle exec rspec spec/lib/gitlab/database/aggregation/click_house/traversal_path_dimension_spec.rb spec/lib/gitlab/database/aggregation/click_house/has_any_filter_spec.rbbundle exec rspec ee/spec/models/analytics/aggregation_engines/ai_usage_events_spec.rbbundle exec rspec ee/spec/requests/api/graphql/analytics/ai_analytics/duo_usage_events_spec.rb
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.