Loading
feat: add content_omitted flag to skip Line/Before/After bytes in search results
What does this MR do and why?
Implements the indexer side of gitlab#606172 (closed) — "Reduce Zoekt aggregation payload by skipping content bytes".
Background
The aggregation-mode Zoekt search path in Rails (introduced in gitlab!244556 (merged)) sets aggregation_mode: true, which raises max_line_match_results_per_file to 5000 so per-file match counts are untruncated. The Rails tally logic only reads file.Language and file.LineMatches.size — it never reads LineMatches[].Line, .Before, or .After (base64-encoded content bytes). For hit-heavy queries this wastes:
- Bandwidth between the indexer and Rails
- CPU on both sides (base64 encoding here, JSON parsing there)
- Shard I/O to load content that isn't needed
Changes
internal/search/types.go
- Added
ContentOmitted boolfield (json:"content_omitted") to bothSearchRequestandRawSearchRequest.
internal/search/search_request.go
- Propagates
ContentOmittedfromRawSearchRequest→SearchRequestinconvertV2.
internal/search/grpc.go
- In
DoSearch: whenContentOmittedis true, forcesNumContextLines = 0in the upstream gRPCSearchOptionsso the Zoekt shard does not load context lines from disk. - In
convertGrpcFiles/convertGrpcLineMatches: whencontentOmittedis true,Line,Before, andAfterare left nil on eachLineMatch. All other fields (LineNumber,LineFragments,LineStart,LineEnd,Score,Language, etc.) remain fully populated.
What is preserved (Rails tally requirements)
file.Language✅ file.LineMatchescount (untruncated)✅ LineMatches[].LineNumber✅ LineMatches[].LineFragments✅ - All counting/truncation behaviour (
max_line_match_results_per_file,max_line_match_window, etc.)✅
What is omitted when content_omitted: true
LineMatches[].Line(the matched line bytes)LineMatches[].Before(context line before)LineMatches[].After(context line after)- Context lines are also not requested from the shard (
NumContextLines = 0)
Edited by Ravi Kumar