ActiveContext: add apply_embeddings_by_root_namespace preprocessor

What does this MR do and why?

Problem and Solution overview

For SaaS (GitLab.com), AIGW requests require the x-gitlab-root-namespace-id header for billing and compliance tracking. However, the semantic search indexing process does not set this header when invoking embeddings requests. This is because indexing is an asynchronous and batched process without a user and is namespace-agnostic.

In order to resolve this, we need to get the root_namespace_id from the projects being indexed, then set that root_namespace_id in the header. The implementation involves several steps:

  • Add a preprocessor to set the root_namespace_id of each reference based on the project_id (Code collection routing) - !246906 (merged)
  • Update the Gitlab::Llm::Embeddings::* classes to accept an optional root_namespace_id, to be used for the AIGW request header - This MR - !246851 (merged)
  • Update the embeddings preprocessor to group embeddings generation by root_namespace_id. This change has to be further split into 2 MRs because of the size:
    • Introduce an apply_embeddings_by_root_namespace preprocessor - This MR
    • Use the apply_embeddings_by_root_namespace in the References::Code class - !248895 (merged)

For details, see proposal.

Updates in this MR

  • Update the ActiveContext::EmbeddingModel#generate_embeddings
    • accept a root_namespace_id parameter
    • pass the root_namespace_id parameter to the embeddings llm_class
  • Introduce a new preprocessor for embeddings by namespace: apply_embeddings_by_root_namespace
    • groups the references by root_namespace_id before processing each group

Note: at this point, we are NOT YET using the new preprocessor for the Semantic Search Indexing pipeline. This MR is simply to introduce the capability.

References

Issue: https://gitlab.com/gitlab-org/gitlab/-/work_items/606107+

Screenshots or screen recordings

N/A

How to set up and validate locally

Since we are not yet introducing the apply_embeddings_by_root_namespace to the Semantic Search Indexing pipeline, we only need to test that the current processing logic is still working as expected:

  1. Add Code collection documents to the processing queue. You may do this by indexing a new project (see instructions) or simply queue an already indexed document to the bulk processing queue:

    ::Ai::ActiveContext::Collections::Code.track_refs!(
      routing: "1000000",
      hashes: [
        "f46dddc837ef559dd22e6875f947bf632c12920fcd2486f7398d66f38a62eba1",
        "2fcd166006b3b99f673672d773b917a075f2e3808ad482b8a5829b16f7f91304"
      ]
    )
  2. Run the bulk processor:

    ::Ai::ActiveContext::BulkProcessWorker.new.perform("Ai::ActiveContext::Queues::Code", 0)
  3. Check the logs:

    Rails:

    # On the GitLab root directory
    tail -f log/active_context.log
    {"severity":"INFO","time":"2026-08-05T04:01:36.139Z","class_name":"ActiveContext::BulkProcessQueue","queue":"Ai::ActiveContext::Queues::Code","message":"bulk_indexing_start","meta.indexing.redis_set":"ai_activecontext_queues:{code}:0:zset","meta.indexing.refs_count":2,"meta.indexing.first_score":1.0,"meta.indexing.last_score":2.0}
    {"severity":"INFO","time":"2026-08-05T04:01:36.157Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","message":"ActiveContext client search completed","collection":"Ai::ActiveContext::Collections::Code","duration_s":0.012,"result_count":2}
    {"severity":"INFO","time":"2026-08-05T04:01:36.212Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","message":"generate embeddings","model":"gitlab_managed__text_embedding_005_vertex","status":"start","class_name":"ActiveContext::EmbeddingModel"}
    {"severity":"INFO","time":"2026-08-05T04:01:48.287Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","message":"generate embeddings","model":"gitlab_managed__text_embedding_005_vertex","status":"done","class_name":"ActiveContext::EmbeddingModel"}
    
    # log indicating errors_count=0
    {"severity":"INFO","time":"2026-08-05T04:01:48.363Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","message":"bulk_submitted","meta.indexing.bulk_count":2,"meta.indexing.errors_count":0}
    
    {"severity":"INFO","time":"2026-08-05T04:01:48.366Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","class_name":"ActiveContext::BulkProcessQueue","message":"bulk_indexer_flushed","meta.indexing.flushing_duration_s":0.0741660000057891}
    
    # log indicating that 2 documents were processed (see meta.indexing.refs_count) 
    {"severity":"INFO","time":"2026-08-05T04:01:48.367Z","correlation_id":"a6f698a8ad6bdf4d59497f1e4b2e8a56","class_name":"ActiveContext::BulkProcessQueue","message":"bulk_indexing_end","meta.indexing.redis_set":"ai_activecontext_queues:{code}:0:zset","meta.indexing.refs_count":2,"meta.indexing.first_score":1.0,"meta.indexing.last_score":2.0,"meta.indexing.failures_count":0,"meta.indexing.retryable_count":0,"meta.indexing.bulk_execution_duration_s":12.227619000012055}

    AIGW:

    # on the GDK root directory
    tail -f log/gitlab-ai-gateway/gateway_debug.log
    
    # this log entry indicates the embeddings request was successful
    # at this point, the `gitlab_root_namespace_id` is still set to `None`
    2026-08-05T04:01:48.269649Z [info     ] 172.16.123.1:55004 - "POST /v1/embeddings/code_embeddings/index HTTP/1.1" 200 [api.access] 
      client_ip=172.16.123.1 client_port=55004 content_type=application/json correlation_id=a6f698a8ad6bdf4d59497f1e4b2e8a56 
      cpu_s=0.12743199999999888 duration_request=0.03701519966125488 duration_s=11.900940666004317 
      enabled-instance-verbose-ai-logs=False enabled_feature_flags= first_chunk_duration_s=11.900914333004039 
      gitlab_feature_enabled_by_namespace_ids= gitlab_feature_enablement_type= gitlab_global_user_id=None 
      gitlab_host_name=gdk.test gitlab_instance_id=7b8ebdcb-fa06-44fe-bf65-2e9dfafe670a gitlab_language_server_version=None 
      gitlab_organization_id=None gitlab_realm=self-managed gitlab_root_namespace_id=None gitlab_saas_duo_pro_namespace_ids=None 
      gitlab_version=19.3.0 http_version=1.1 is_gitlab_team_member=false meta.feature_category=global_search method=POST 
      path=/v1/embeddings/code_embeddings/index request_arrived_at=2026-08-05T04:01:36.368524+00:00 response_start_duration_s=11.900883207999868 
      stage=main status_code=200 type=mlops url=http://gdk.test:5052/v1/embeddings/code_embeddings/index user_agent=Ruby

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #606107

Edited by Pam Artiaga

Merge request reports

Loading
Loading