Aggregation engine queries cannot set per-query ClickHouse settings such as max_execution_time
Problem
Analytics aggregation queries run through Gitlab::Database::Aggregation::ClickHouse::AggregationResult, which calls ClickHouse::Client.select(query, :main). ClickHouse::Client.select(query, database, configuration) takes no per-query settings argument, so the only settings applied come from the per-database variables in config/click_house.yml, which apply to every ClickHouse query in the application, not just aggregation queries. There is currently no way to put a max_execution_time or max_rows_to_read ceiling on aggregation queries specifically.
This came up during review of !256398 (merged), which adds exact_not_match exclusion filters to 8 ClickHouse-backed aggregation engines. Exclusion filters compile to NOT IN, which cannot use the ClickHouse sort key to skip granules the way IN can, so an exclusion query reads the whole scoped range. Aggregation filters are all registered required: false and there is no default time range, so a broad query has no upper bound on the amount of work it can trigger.
Evidence
- A
grepformax_execution_timeandmax_rows_to_readacrosslib,app,ee, andconfigreturns nothing: there is no existing per-query settings mechanism for aggregation queries. ClickHouse::Client::Database#build_custom_uri(extra_variables:)already exists and is public, andinsert_csvalready uses it. This suggests a per-query settings path may be a smaller addition than it first appears (pointer raised by @terrichu during review).- Local
EXPLAIN indexes = 1numbers taken during that review showedNOT INselecting 306 of 306 granules where the matchingINselected 1, on 2.5 million rows. That benchmark placed all rows under a singletraversal_path, which is the first sort-key column, so the path prefix pruned nothing. In production the prefix narrows first and the exclusion only scans within it, so 306/306 is a worst case, not necessarily the expected case. The number that would actually settle this is the row count under the largest namespace inai_usage_events,agent_platform_sessions, andai_code_suggestions. Per @terrichu, pulling that figure is on the Optimize/analytics side.
Proposed direction
Decide whether aggregation queries should carry a query-level ceiling and, if so, whether it is expressed as a SETTINGS clause on the generated SQL or as a per-query option threaded through the ClickHouse client. Either approach changes behavior for every aggregation query, including the existing positive filters, so it needs its own performance review before landing.