Bug: No validation of embedding count against input count in apply_embeddings — silent nil corruption possible
Summary
ActiveContext::Preprocessors::Embeddings#apply_embeddings does not validate that the number of embeddings returned by the AI Gateway matches the number of input documents. If the API returns fewer embeddings than expected for any reason, the mismatch is silently ignored and affected documents get nil written as their embedding value in OpenSearch — corrupting the search index with no error, log, or retry.
Root cause
gems/gitlab-active-context/lib/active_context/preprocessors/embeddings.rb lines 47–61:
model_groups.each_value do |items|
models = items.first[:models]
contents = items.map { |item| item[:doc][content_field] } # N items
embeddings_by_model = generate_embeddings_for_each_model(models: models, contents: contents)
items.each.with_index do |item, index|
models.each do |model|
item[:doc][model.field] = embeddings_by_model[model.field][index] # nil if index >= returned count
end
end
endembeddings_by_model[model.field] is the raw array returned from Gitlab::Llm::Embeddings::CodeEmbeddings#execute → response.embeddings → parsed_response['predictions'].pluck('embedding'). If the API returns fewer predictions than the number of contents sent (e.g. 3 returned for 5 sent), indexing beyond the array returns nil with no exception raised.
Gitlab::Llm::Embeddings::Response#embeddings similarly has no count check:
def embeddings
return unless success?
(parsed_response['predictions'] || []).pluck('embedding') # no validation against input count
endImpact
- Documents that receive
nilembeddings are written to OpenSearch with a null vector field - Subsequent KNN searches against those documents produce garbage results silently
- There is no retry, no error log, and no way to detect which documents were corrupted short of querying OpenSearch directly
- The bug is in the gem's preprocessor and affects both the
BulkProcessWorkerproduction pipeline and any direct caller ofapply_embeddings
When could this happen?
The production path goes through AIGW which normalizes Vertex AI responses, so a count mismatch is not a common occurrence today. However there is no contract enforcement anywhere in the stack, making this a fragile assumption. Possible triggers include:
- An AIGW or upstream Vertex AI partial response under load
- A future change to batching logic that introduces an off-by-one
- A self-hosted AI Gateway implementation that returns a different response shape
Expected behaviour
Before writing embeddings back to documents, validate that the returned count matches the input count. If mismatched, raise an error (or log and skip the batch) rather than silently writing nil:
embeddings = embeddings_by_model[model.field]
if embeddings.length != items.length
raise IndexingError, "Embedding count mismatch: expected #{items.length}, got #{embeddings.length}"
end
items.each.with_index do |item, index|
item[:doc][model.field] = embeddings[index]
endA similar guard should be added in Gitlab::Llm::Embeddings::Response#embeddings or CodeEmbeddings#execute to detect and surface mismatches at the API boundary before they propagate to the indexing layer.