Raise the ClickHouse aggregation page size cap to 250

What does this MR do and why?

Aggregation engine GraphQL fields (the aggregated connection under each engine's ...AggregationScope type, for example DuoWorkflowsAggregationScope) were capped at the schema-wide default_max_page_size of 100 rows per page. Every page re-runs the whole ClickHouse aggregation and then skips earlier rows with OFFSET. Clients that walk every page, such as GLQL charts on dashboards like DAP Impact, repeat that full aggregation once per 100 rows. This MR raises the page cap for ClickHouse aggregation engines to 250, so fewer requests are needed to page through the same data.

  • Gitlab::Database::Aggregation::Engine.max_page_size returns nil, so it keeps the schema default.
  • Gitlab::Database::Aggregation::ClickHouse::Engine.max_page_size returns 250.
  • Resolvers::Analytics::Aggregation::AggregationFieldResolver.build sets max_page_size on the field when the engine returns one.
  • All 8 mounted engines are ClickHouse engines, all mounted from EE (ee/app/graphql/types/analytics/analytics_type.rb), so every aggregation field now allows up to 250 rows.
  • The larger cap is behind the larger_clickhouse_aggregation_pages gitlab_com_derisk flag, with the current user as the actor. The field cap is fixed when the schema loads, so AggregationConnection#max_page_size checks the flag on each request. With the flag off, every engine keeps the schema default of 100.

One behaviour to flag for reviewers: with the flag on, clients that omit first get up to 250 rows instead of 100, because graphql-ruby falls back to the max page size when no default page size is set. I did not set default_page_size on purpose. Some frontend clients that omit first, such as the Duo feature retention query that pages with after only, benefit from fewer, larger pages. I also could not reference GitlabSchema.default_max_page_size inside the resolver builder, because that code runs while the schema is still being defined and it breaks schema loading.

250 is a starting point for now and can be tuned later. Follow-up: !257576 (merged) makes GLQL request pages of this size.

References

Parent work item: https://gitlab.com/gitlab-org/gitlab/-/work_items/630507 (DAP Impact V1 - Performance improvements), item "Bigger aggregation pages".

Feature flag rollout: #630765

Screenshots or screen recordings

Not applicable, this MR has no UI change.

How to set up and validate locally

  1. Check out this branch.
  2. In bin/rails console, enable the flag with Feature.enable(:larger_clickhouse_aggregation_pages), then run GitlabSchema.types['DuoWorkflowsAggregationScope'].fields['aggregated'].max_page_size and expect 250.
  3. In GraphiQL at /-/graphql-explorer, run an aggregation query with aggregated(first: 250) against a group with ClickHouse data. You should see up to 250 nodes returned, where it used to stop at 100.

Locally I also ran a rails runner script that printed max_page_size=250 for all 8 *AggregationScope.aggregated fields and default_max_page_size=100 for the schema. A new spec, spec/graphql/resolvers/analytics/aggregation/aggregation_field_resolver_spec.rb, checks that a ClickHouse engine gets 250 and an ActiveRecord engine keeps nil. New cases in spec/lib/gitlab/database/aggregation/graphql/aggregation_connection_spec.rb check that the connection uses the field cap when the flag is on for the current user, and falls back to 100 when it is off.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

🤖 Generated with Claude Code

Edited by Brandon Labuschagne

Merge request reports

Loading
Loading