feat(search): add language aggregation to the v2 search API

What

Adds an opt-in aggregations: ["language"] key to the v2 search API. Requests that ask for it get per-language match counts tallied by the indexer across the whole result stream, rather than the client tallying whatever happened to fit in the returned payload.

Why

Language buckets were tallied from the returned payload, which is capped at max_line_match_window (5000 line matches). Once that window filled, whole files were dropped, so a low-ranked language reported a fraction of its real match count — while ticking that language's filter returned the true count from Zoekt's own stats. The sidebar and the result header disagreed by orders of magnitude.

The root cause is that one number bounded two unrelated things: payload size and scan depth. max_line_match_window is forwarded to Zoekt as TotalMaxMatchCount, which stops it searching further shards, so the counts could never see past the point where the returned matches stopped.

How

  • aggregations: ["language"] opts in. aggregation_match_budget (default 100k, ceiling 1M) bounds scan depth separately from payload size — that separation is the fix.
  • Matching files are tallied as they stream in and their LineMatches dropped, so the response is O(languages) rather than O(matches).
  • Cross-node replica dedup is preserved by retaining a lightweight record per file instead of the full FileMatch.
  • The response carries Result.Aggregations, emitted only when aggregations were requested.
  • Unknown dimension names are ignored rather than rejected, so a newer client can ask for dimensions this indexer does not implement.

Compatibility — no version gate on either side

  • Old client, new indexer: nothing sends aggregations, the branch is never entered, and Result.Aggregations is never emitted. Existing traffic is unchanged.
  • New client, old indexer: the key is ignored and no Aggregations appears, which is the client's signal to fall back to tallying the payload.
  • Mixed fleet during rollout: only the entry node interprets aggregations. The gRPC leg to other nodes is a plain StreamSearch with standard SearchOptions, so downstream nodes need no particular version.

Either deploy order is safe.

Performance

Measured against a two-node local index. The aggregation-mode call, comparing a client-side tally against this one:

latency response client-side JSON parse languages tallied / true
client tally (5k window) 13.6 ms 3,893,546 B ~5.5 ms 20 4,875 / 10,850
indexer aggregation 7.6 ms 572 B ~0 26 10,850

The payload drops by roughly 6,800×, and it is O(languages) — flat at 484–620 bytes across every query tested, where the client-tally payload grows with the corpus. Faster overall despite scanning 20× deeper, because serialising megabytes of line matches costs more than scanning further.

A follow-up commit (46e1a3d) shrinks what the tally retains. One record is held per matched file for the life of the request, and unlike the file path — where a dedup key rides alongside kilobytes of match content — here the matches are dropped and the record is the memory. Over 33,000 files across 20 languages:

retained allocations time
before 7,400,080 B 132,002 5.12 ms
after 1,058,568 B 6 1.92 ms

The key is only ever compared, so it became a 64-bit hash streamed through a digest on the stack; language names are interned. Only the aggregation path changes — the file path keeps its readable string key, where the set is bounded by the payload window and an opaque hash would cost clarity for nothing.

Known limitation: truncated counts are not lower bounds

Worth knowing before this is relied on. When the budget runs out (Truncated: true), the counts are not proportional-but-small — they are biased by shard search and arrival order, and the ranking can be wrong at the top.

Measured on one node, raising the budget for a single query, showing each language's share of that run's total:

budget tallied HTML JavaScript Markdown C
5,000 5,088 0.3% 34.1% 44.8% 0.0%
50,000 50,051 2.7% 65.4% 12.5% 0.0%
complete 233,820 34.3% 19.3% 3.5% 12.6%

At a 5,000 budget the largest language in the corpus shows as 0.3%, and on a second query the true top language was absent from the top six entirely.

Two cutoffs cause it, neither ordered by relevance or match density. Zoekt (search/shards.go) dispatches shards sorted by RawConfig["priority"] then repository name, and stop() halts dispatch of further shards once accumulated matches exceed TotalMaxMatchCount — whole shards are skipped, not "the last N matches". This indexer never sets RawConfig, so all shards tie at priority 0. Separately, handleGrpcSearchStream stops consuming at the budget, discarding chunks already in flight.

This is not a regression — a client-side tally of the first 5,000 payload matches has exactly the same bias; it had simply never been measured. This change is strictly better because it can complete, and when it does the counts are exact. But a bigger budget does not buy proportionally better counts: only completing the scan does.

The consequence for clients: treat Truncated: true as "unreliable", not "low". Whether to render buckets in that state is a product decision worth making deliberately, since a confidently wrong ranking may be worse than no sidebar. For contrast, Elasticsearch truncates terms aggregations by rank, so it loses the tail and never the largest bucket.

Also removes the content_omitted flag

86a163e reverts !989 (merged). content_omitted stripped Line/Before/After from LineMatches so the aggregation-mode search would stop shipping content bytes the client never read. This change supersedes it: an aggregating response carries no files at all rather than a stripped-down copy of them, and the aggregating path already forces NumContextLines = 0 on its own, so the shard does not load context lines either. The flag has nothing left to do.

No caller is affected. content_omitted appears nowhere in GitLab Rails — not on master, not on the client-side branch for this feature — so no request in existence sets it, and at false it was always a no-op. Nothing marshals SearchRequest outbound either, so dropping the field changes what the indexer accepts, never what it emits. Both deploy orders stay safe for the same reasons as above.

It reverts d3f2cad exactly. convertGrpcFiles/convertGrpcLineMatches and both test files it touched are byte-identical to their pre-d3f2cad state, verified with git diff d3f2cad^. search_request_test.go held only its test, so the file goes with it. The v1.18.0 changelog entry stays, since that release did ship the flag.

For reviewers

Two commits deserve more than a skim. 46e1a3d is not purely additive. combineResults is on the hot path for every search, and the dedup-map allocation there was restructured. It is behaviour-neutral, but it deserves an eye rather than being waved through as aggregation-only code.

Verified by building the pre-change binary in a separate worktree and diffing a non-aggregating search against it: identical FileCount, MatchCount, file set and per-file match counts. Note that a truncated response is not byte-stable across runs on any build, because max_line_match_window cuts wherever the Nth match lands and that depends on which node's chunks arrive first — so the comparison was made with the window raised above the result size.

The aggregating response was separately confirmed byte-identical (same SHA-256) before and after the performance commit.

86a163e is a revert, but it edits convertGrpcLineMatches, which every search goes through — worth confirming the converters really are back to their pre-!989 form rather than merely equivalent.

The last three commits are readability cleanups with no behaviour change: min for the budget clamp, slices.Contains for the dimension check, and a single indexed lookup in place of a loop over dimensions that could only ever visit one. The existing aggregation tests cover all three unchanged.

Test coverage

merge and newLanguageAggregator at 100%, combineResults at 94.9%, package 91.5%. Tests cover: counts surviving the payload window and per-file cap, the node-side budget, the merge-side budget signalling the fan-out to stop, combineResults acting on that signal, a node reporting its own truncation, timeout handling, response-key presence and absence, unknown dimensions, and cross-node replica dedup.

One test records that identical paths in different repositories are distinct files. That invariant predates this change, but hashing the dedup key makes it invisible to anyone narrowing the key later, so it is now pinned by a failing test rather than by reading the code.

Related: gitlab#392882 (closed)

🤖 Generated with Claude Code

Edited by Ravi Kumar

Merge request reports

Loading
Loading