Semantic Search indexing: apply embeddings by root namespace in SaaS

What does this MR do and why?

Problem and Solution overview

For SaaS (GitLab.com), AIGW requests require the x-gitlab-root-namespace-id header for billing and compliance tracking. However, the semantic search indexing process does not set this header when invoking embeddings requests. This is because indexing is an asynchronous and batched process without a user and is namespace-agnostic.

In order to resolve this, we need to get the root_namespace_id from the projects being indexed, then set that root_namespace_id in the header. The implementation involves several steps:

  • Add a preprocessor to set the root_namespace_id of each reference based on the project_id (Code collection routing) - !246906 (merged)
  • Update the Gitlab::Llm::Embeddings::* classes to accept an optional root_namespace_id, to be used for the AIGW request header - This MR - !246851 (merged)
  • Update the embeddings preprocessor to group embeddings generation by root_namespace_id. This change has to be further split into 2 MRs because of the size:
    • Introduce an apply_embeddings_by_root_namespace preprocessor - !248475 (merged)
    • Use the apply_embeddings_by_root_namespace in the References::Code class - This MR

For details, see proposal.

Updates in this MR

  • use the apply_embeddings_by_root_namespace in the Ai::ActiveContext::References::Code class
    • this is only done in the SaaS instance
    • this is toggled behind a Feature Flag
  • Add logging for both the user_id and root_namespace_id for embeddings request

References

Issue: https://gitlab.com/gitlab-org/gitlab/-/work_items/606107+

Screenshots or screen recordings

N/A

How to set up and validate locally

  1. Simulate SaaS instance by setting environment variable: GITLAB_SIMULATE_SAAS=1

  2. Enable the Feature Flag: Feature.enable(:semantic_search_indexing_set_root_namespace_id)

  3. Add Code collection documents to the processing queue. You may do this by indexing a new project (see instructions) or simply queue an already indexed document to the bulk processing queue:

    ::Ai::ActiveContext::Collections::Code.track_refs!(
      routing: "1000000",
      hashes: [
        "450532041f9569f7aeeb7d3d59f984d5c5e4d2363ed9613135ec265f26eb88ca",
        "8e0d0f24bc2cb6c4bc2a42cd5af55877991acc2319d16c183000748dfe0c83da"
      ]
    )
  4. Run the bulk processor:

    ::Ai::ActiveContext::BulkProcessWorker.new.perform("Ai::ActiveContext::Queues::Code", 0)
  5. Check the logs:

    • Rails active_context.log

      # on the GitLab root directory
      tail -f log/active_context.log
      {"severity":"INFO","time":"2026-08-07T04:44:48.982Z","class_name":"ActiveContext::BulkProcessQueue","queue":"Ai::ActiveContext::Queues::Code","message":"bulk_indexing_start","meta.indexing.redis_set":"ai_activecontext_queues:{code}:0:zset","meta.indexing.refs_count":2,"meta.indexing.first_score":3.0,"meta.indexing.last_score":4.0}
      {"severity":"INFO","time":"2026-08-07T04:44:48.998Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"ActiveContext client search completed","collection":"Ai::ActiveContext::Collections::Code","duration_s":0.011,"result_count":2}
      {"severity":"INFO","time":"2026-08-07T04:44:48.998Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"Resolving root_namespace_id for references","class_name":"Ai::ActiveContext::References::Code","preprocessor":"code_root_namespace_resolver","refs_count":2}
      {"severity":"INFO","time":"2026-08-07T04:44:49.036Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"Resolved root_namespace_id for references","class_name":"Ai::ActiveContext::References::Code","preprocessor":"code_root_namespace_resolver","refs_count":2,"refs_with_root_namespaces_count":2,"unique_root_namespace_ids":[1000000]}
      
      # indicates embeddings being generated
      {"severity":"INFO","time":"2026-08-07T04:44:49.044Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"generate embeddings","model":"gitlab_managed__text_embedding_005_vertex","status":"start","class_name":"ActiveContext::EmbeddingModel"}
      {"severity":"INFO","time":"2026-08-07T04:45:01.056Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"generate embeddings","model":"gitlab_managed__text_embedding_005_vertex","status":"done","class_name":"ActiveContext::EmbeddingModel"}
      
      # indicates that the 2 queued references were processed successfully
      {"severity":"INFO","time":"2026-08-07T04:45:01.121Z","correlation_id":"72381112ad877ce332306394151ad83e","message":"bulk_submitted","meta.indexing.bulk_count":2,"meta.indexing.errors_count":0}
      {"severity":"INFO","time":"2026-08-07T04:45:01.125Z","correlation_id":"72381112ad877ce332306394151ad83e","class_name":"ActiveContext::BulkProcessQueue","message":"bulk_indexer_flushed","meta.indexing.flushing_duration_s":0.06606700003612787}
      {"severity":"INFO","time":"2026-08-07T04:45:01.125Z","correlation_id":"72381112ad877ce332306394151ad83e","class_name":"ActiveContext::BulkProcessQueue","message":"bulk_indexing_end","meta.indexing.redis_set":"ai_activecontext_queues:{code}:0:zset","meta.indexing.refs_count":2,"meta.indexing.first_score":3.0,"meta.indexing.last_score":4.0,"meta.indexing.failures_count":0,"meta.indexing.retryable_count":0,"meta.indexing.bulk_execution_duration_s":12.143157999962568}
    • Rails llm.log

      # on the GitLab root directory
      tail -f log/llm.log
      
      # Note that this log entry now has a `root_namespace_id=1000000`
      {"severity":"INFO","time":"2026-08-07T04:44:49.050Z","correlation_id":"72381112ad877ce332306394151ad83e","unit_primitive":"generate_embeddings_codebase","url":"http://gdk.test:5052/v1/embeddings/code_embeddings/index","user_id":null,"root_namespace_id":1000000,"params":"{\"model_metadata\":{\"provider\":\"gitlab\",\"identifier\":\"text_embedding_005_vertex\"}}","message":"Performing embeddings request","class":"Gitlab::Llm::Embeddings::Client","ai_event_name":"performing_request","ai_component":"code_embeddings_index"}
      
      {"severity":"INFO","time":"2026-08-07T04:45:01.056Z","correlation_id":"72381112ad877ce332306394151ad83e","unit_primitive":"generate_embeddings_codebase","url":"http://gdk.test:5052/v1/embeddings/code_embeddings/index","message":"Received embeddings response","class":"Gitlab::Llm::Embeddings::Client","ai_event_name":"response_received","ai_component":"code_embeddings_index"}
    • AIGW debug log

      # on the GDK root directory
      tail -f log/gitlab-ai-gateway/gateway_debug.log
      2026-08-07T04:45:01.050026Z [info     ] 172.16.123.1:56666 - "POST /v1/embeddings/code_embeddings/index HTTP/1.1" 200 [api.access] 
        client_ip=172.16.123.1 client_port=56666 content_type=application/json correlation_id=72381112ad877ce332306394151ad83e 
        cpu_s=0.13243599999998423 duration_request=0.02727985382080078 duration_s=11.903599541983567 enabled-instance-verbose-ai-logs=False 
        enabled_feature_flags= first_chunk_duration_s=11.903574958996614 gitlab_feature_enabled_by_namespace_ids= gitlab_feature_enablement_type= 
        gitlab_global_user_id=None gitlab_host_name=gdk.test gitlab_instance_id=7b8ebdcb-fa06-44fe-bf65-2e9dfafe670a gitlab_language_server_version=None 
        gitlab_organization_id=None gitlab_realm=saas 
        gitlab_root_namespace_id=1000000 # this header field now has a value
        gitlab_saas_duo_pro_namespace_ids=None 
        gitlab_version=19.3.0 http_version=1.1 is_gitlab_team_member=false meta.feature_category=global_search method=POST 
        path=/v1/embeddings/code_embeddings/index request_arrived_at=2026-08-07T04:44:49.146269+00:00 
        response_start_duration_s=11.90354870900046 stage=main status_code=200 type=mlops 
        url=http://gdk.test:5052/v1/embeddings/code_embeddings/index user_agent=Ruby
  6. As a final validation step, you can also check your Elasticsearch index and verify that the processed documents now has embeddings

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #606107

Edited by Pam Artiaga

Merge request reports

Loading
Loading