Use indexer-side language counts for Zoekt code search

Rails half of indexer-side language aggregation. Pairs with and depends on gitlab-zoekt-indexer!999 (merged) — without an indexer that understands aggregations, the language sidebar renders empty.

Behind zoekt_language_aggregations, which is default-off and has never been enabled anywhere.

What

The language sidebar for Zoekt code search now uses counts the indexer tallied across the whole result stream, instead of counting whatever fitted in the returned payload. It also gains a row for files whose language could not be detected.

How

Use the indexer's counts, with no client-side fallback. The payload tally is removed rather than kept as a backstop. It only ever ran behind a flag that has never been enabled, so there is no shipped behaviour to preserve: an indexer that cannot aggregate now yields an empty sidebar, which is exactly what flag-off already produces. Keeping it would have meant two paths that disagree on the same numbers, with nothing to tell a user which one they are looking at.

That also retires the machinery which existed only to reconcile the two — the exact: flag on AggregationCache#write and the per-file cap lift in Params. The cap could not have affected the counts in any case: the indexer bounds aggregation scan depth with its own aggregation_match_budget and drops LineMatches as it tallies. Response#aggregations keeps its nil-versus-empty distinction, which turned out to have a second use: it now picks the cache TTL rather than choosing between two tallying paths.

Either deploy order remains safe, but the flag must not be enabled until the indexer change is deployed. There is no version gate, so that ordering is a rollout constraint rather than something the code enforces.

Unclassified files get a row. The indexer reports them under an empty language and they were dropped, because a bucket you cannot filter by is a checkbox that contradicts its own count. On a broad query that silently hid about a fifth of the matches.

Zoekt has no atom for "no detected language" — lang: looks the name up in the shard's language map, and a name that is absent matches nothing rather than matching the unclassified files. The filter is therefore the complement of every language in the tally:

def (-lang:Batchfile -lang:C -lang:Go -lang:HTML … -lang:YAML)

Verified against a local index: a query with 10,850 matches of which 969 are unclassified returns exactly 969, and 1,012 when Unknown is ticked alongside Ruby (969 + Ruby's 43).

It fails open. The language list comes from a cache that the aggregations request populates, and known_languages now fetches the tally itself on a miss instead of only reading, so ordering between /search and /search/aggregations no longer matters. What remains is a truncated tally: if the indexer ran out of match budget, the tally is a subset of the corpus, known_languages returns [], unknown_atom returns nil, and the Unknown filter is compacted away rather than applied — any other languages the user ticked still apply as normal lang: atoms, so only the Unknown term disappears. The review rejected failing closed here, because truncation also drops the Unknown row from the sidebar, so an empty page would leave the user with zero results and nothing to untick. Dropping the filter instead means the user still sees results, just not the filtered set they asked for, with no signal that the Unknown selection was ignored — a known, accepted gap tracked in #623415 (closed). The follow-up has since been decided at the sidebar level: rather than labelling a truncated tally "approximate", which would misrepresent a biased sample as merely imprecise, the sidebar will be suppressed entirely under Truncated. That suppression is not implemented here, and once it lands, this MR's fail-open behaviour stops being reachable through the sidebar checkbox but remains reachable via a URL-supplied filter, which is why the follow-up also requires telling the user the filter could not be applied.

Cache key. Search::Zoekt::Cache#search_fingerprint now includes known_languages (JSON-serialised, alongside the existing filters key). The same user-visible filter selection can produce a different Zoekt query depending on which languages were in the tally at the time, so without it two different queries would collide on one cached result set.

Presentation. Unknown sorts last regardless of size and is exempted from the display cut, so it does not get dropped when the bucket list is truncated. This MR emits extra: { 'deemphasized' => true } on that bucket — a payload flag for the frontend to consume — but does not render anything itself; the rendering (greying out the row) is in a separate MR, !252040 (merged). The earlier plan to keep a second flag, filterable: false, for rows that "genuinely cannot be selected" turned out to be unnecessary: nothing in the codebase, Ruby or JS, ever emitted it. It was built in anticipation of Unknown being unexpressible as a lang: atom, which the complement filter disproved. It has been deleted in the frontend MR, so buckets now carry one flag instead of two.

Bucket cap. Buckets are capped at the same ceiling Advanced Search puts on its own language aggregation, so the two backends cannot render sidebars of different lengths for the same corpus. The cap is display-only — the Unknown filter still excludes every language in the tally, not just the visible ones, because excluding only the visible ones would report the languages past the cut as unclassified. That is covered by a test rather than a comment.

References

🤖 Generated with Claude Code

Edited by Ravi Kumar

Merge request reports

Loading
Loading