Aggregation engine queries cannot set per-query ClickHouse settings such as max_execution_time

Problem

Analytics aggregation queries run through Gitlab::Database::Aggregation::ClickHouse::AggregationResult, which calls ClickHouse::Client.select(query, :main). ClickHouse::Client.select(query, database, configuration) takes no per-query settings argument, so the only settings applied come from the per-database variables in config/click_house.yml, which apply to every ClickHouse query in the application, not just aggregation queries. There is currently no way to put a max_execution_time or max_rows_to_read ceiling on aggregation queries specifically.

This came up during review of !256398 (merged), which adds exact_not_match exclusion filters to 8 ClickHouse-backed aggregation engines. Exclusion filters compile to NOT IN, which cannot use the ClickHouse sort key to skip granules the way IN can, so an exclusion query reads the whole scoped range. Aggregation filters are all registered required: false and there is no default time range, so a broad query has no upper bound on the amount of work it can trigger.

Evidence

  • A grep for max_execution_time and max_rows_to_read across lib, app, ee, and config returns nothing: there is no existing per-query settings mechanism for aggregation queries.
  • ClickHouse::Client::Database#build_custom_uri(extra_variables:) already exists and is public, and insert_csv already uses it. This suggests a per-query settings path may be a smaller addition than it first appears (pointer raised by @terrichu during review).
  • Local EXPLAIN indexes = 1 numbers taken during that review showed NOT IN selecting 306 of 306 granules where the matching IN selected 1, on 2.5 million rows. That benchmark placed all rows under a single traversal_path, which is the first sort-key column, so the path prefix pruned nothing. In production the prefix narrows first and the exclusion only scans within it, so 306/306 is a worst case, not necessarily the expected case. The number that would actually settle this is the row count under the largest namespace in ai_usage_events, agent_platform_sessions, and ai_code_suggestions. Per @terrichu, pulling that figure is on the Optimize/analytics side.

Proposed direction

Decide whether aggregation queries should carry a query-level ceiling and, if so, whether it is expressed as a SETTINGS clause on the generated SQL or as a per-query option threaded through the ClickHouse client. Either approach changes behavior for every aggregation query, including the existing positive filters, so it needs its own performance review before landing.