Bug: No validation of embedding count against input count in apply_embeddings — silent nil corruption possible

Summary

ActiveContext::Preprocessors::Embeddings#apply_embeddings does not validate that the number of embeddings returned by the AI Gateway matches the number of input documents. If the API returns fewer embeddings than expected for any reason, the mismatch is silently ignored and affected documents get nil written as their embedding value in OpenSearch — corrupting the search index with no error, log, or retry.

Root cause

gems/gitlab-active-context/lib/active_context/preprocessors/embeddings.rb lines 47–61:

model_groups.each_value do |items|
  models = items.first[:models]
  contents = items.map { |item| item[:doc][content_field] }  # N items

  embeddings_by_model = generate_embeddings_for_each_model(models: models, contents: contents)

  items.each.with_index do |item, index|
    models.each do |model|
      item[:doc][model.field] = embeddings_by_model[model.field][index]  # nil if index >= returned count
    end
  end
end

embeddings_by_model[model.field] is the raw array returned from Gitlab::Llm::Embeddings::CodeEmbeddings#execute → response.embeddings → parsed_response['predictions'].pluck('embedding'). If the API returns fewer predictions than the number of contents sent (e.g. 3 returned for 5 sent), indexing beyond the array returns nil with no exception raised.

Gitlab::Llm::Embeddings::Response#embeddings similarly has no count check:

def embeddings
  return unless success?
  (parsed_response['predictions'] || []).pluck('embedding')  # no validation against input count
end

Impact

  • Documents that receive nil embeddings are written to OpenSearch with a null vector field
  • Subsequent KNN searches against those documents produce garbage results silently
  • There is no retry, no error log, and no way to detect which documents were corrupted short of querying OpenSearch directly
  • The bug is in the gem's preprocessor and affects both the BulkProcessWorker production pipeline and any direct caller of apply_embeddings

When could this happen?

The production path goes through AIGW which normalizes Vertex AI responses, so a count mismatch is not a common occurrence today. However there is no contract enforcement anywhere in the stack, making this a fragile assumption. Possible triggers include:

  • An AIGW or upstream Vertex AI partial response under load
  • A future change to batching logic that introduces an off-by-one
  • A self-hosted AI Gateway implementation that returns a different response shape

Expected behaviour

Before writing embeddings back to documents, validate that the returned count matches the input count. If mismatched, raise an error (or log and skip the batch) rather than silently writing nil:

embeddings = embeddings_by_model[model.field]

if embeddings.length != items.length
  raise IndexingError, "Embedding count mismatch: expected #{items.length}, got #{embeddings.length}"
end

items.each.with_index do |item, index|
  item[:doc][model.field] = embeddings[index]
end

A similar guard should be added in Gitlab::Llm::Embeddings::Response#embeddings or CodeEmbeddings#execute to detect and surface mismatches at the API boundary before they propagate to the indexing layer.

Edited by 🤖 GitLab Bot 🤖