Limit throughput of Code Embeddings Indexing pipeline in SM instances

What does this MR do and why?

We are planning to make the rate limit rule stricter for the AIGW embeddings endpoint for requests coming from SM instances. This stricter rule could result in more rate limit errors, see https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085#note_3211312015. The increased rate limit errors need to be handled in Rails, by:

  1. Make the pipeline more robust so refs with rate limit errors can be retried indefinitely (#595433 (closed))

  2. Slow down the bulk processing queue for SM instances with Gitlab-operated code embeddings models - THIS MR

    Note: This is not the ultimate solution for rate limiting (that will be addressed on the infrastructure level on the AIGW endpoint), but once we do apply stricter rate limiting to the AIGW endpoints, we can pre-emptively reduce the number of 429 errors if we slow down bulk processing.

Implemented solutions in this MR

  • Limit queue throughputs - Set maximum values to the number_of_shards and shard_limit
  • Prevent re-enqueuing - make sure the bulk processor does not re-enqueue

References

Screenshots or screen recordings

N/A

How to set up and validate locally

Setup

  1. Set the queue throughput configuration:

    ::Ai::ActiveContext::Collections::Code.collection_record.update_options!(
      queue_shard_count: 5, # anything higher than 1
      queue_shard_limit: 10_000 # 1000 (default) or higher
    )
  2. Set ActiveContext::Config.re_enqueue_indexing_workers? to true. You can override this in gems/gitlab-active-context/lib/active_context/config.rb

Test in Self-Managed instance

  1. Set environment variable GITLAB_SIMULATE_SAAS=0

  2. Test queue values:

    Ai::ActiveContext::Queues::Code.limit_throughput?
    => true
    
    Ai::ActiveContext::Queues::Code.number_of_shards
    => 1
    
    Ai::ActiveContext::Queues::Code.shard_limit
    => 600
  3. Test the ref tracking:

    # track refs
    ::Ai::ActiveContext::Collections::Code.track_refs!(routing: "1", hashes: [
      "532997495b421186c8868e308b6fd1def6c9988beb76cb91508ea8738700aceb",
      "1311ee1cdd722cce9ac24a16c85212df40b69177b64fa353bbf4a205e6f25016",
      "9f3f66bd3e976a8b50d2cc94ff3bd1d0fd77ac5945f17ca6cf9e55e8925009f4",
      "56c38a4c7df12c221becfa62c410d19746982a9117abc6f8b732d384e100ee2c",
      "29d4bc8ed23583e6d481b18b62f6022f42b9f3df402ec0d2f6b7afb9302d2eeb",
      "6993a86dbb5ea3979cfbc10b6cd4b269ec4702ae8ce742324ceee581617367e6",
      "90ea0421f39ef710ba28f379151e9e87f6c1ceb915aa468b9b5eb732768bc4d4",
      "1229f89f0cc02ec87f16f09649840783786ca4d415d4948322c201b2012a418c",
      "c4087b44b8b5b3030da33a1c5a7490ed02b536202b349b15155308dfcae632a4",
      "8f38deab412aee170b4a81b36c6626b85c7d420cdc1e710e4358d52fb4e480e3"
    ])
    
    # all refs are queued to 1 shard
    ActiveContext::Queues.all_queued_items
    => {"ai_activecontext_queues:{code}:0:zset"=>
    ["Ai::ActiveContext::References::Code|2|1|532997495b421186c8868e308b6fd1def6c9988beb76cb91508ea8738700aceb",
    "Ai::ActiveContext::References::Code|2|1|1311ee1cdd722cce9ac24a16c85212df40b69177b64fa353bbf4a205e6f25016",
    "Ai::ActiveContext::References::Code|2|1|9f3f66bd3e976a8b50d2cc94ff3bd1d0fd77ac5945f17ca6cf9e55e8925009f4",
    "Ai::ActiveContext::References::Code|2|1|56c38a4c7df12c221becfa62c410d19746982a9117abc6f8b732d384e100ee2c",
    "Ai::ActiveContext::References::Code|2|1|29d4bc8ed23583e6d481b18b62f6022f42b9f3df402ec0d2f6b7afb9302d2eeb",
    "Ai::ActiveContext::References::Code|2|1|6993a86dbb5ea3979cfbc10b6cd4b269ec4702ae8ce742324ceee581617367e6",
    "Ai::ActiveContext::References::Code|2|1|90ea0421f39ef710ba28f379151e9e87f6c1ceb915aa468b9b5eb732768bc4d4",
    "Ai::ActiveContext::References::Code|2|1|1229f89f0cc02ec87f16f09649840783786ca4d415d4948322c201b2012a418c",
    "Ai::ActiveContext::References::Code|2|1|c4087b44b8b5b3030da33a1c5a7490ed02b536202b349b15155308dfcae632a4",
    "Ai::ActiveContext::References::Code|2|1|8f38deab412aee170b4a81b36c6626b85c7d420cdc1e710e4358d52fb4e480e3"]}
  4. Test re-enqueue

    • in order to test locally, we need to set a small shard limit.

      # this will be prioritized over the calculated shard limit since it's smaller
      ::Ai::ActiveContext::Collections::Code.collection_record.update_options!(queue_shard_limit: 2)
    • Process the refs queued in Step 4:

      ::Ai::ActiveContext::BulkProcessWorker.new.perform("Ai::ActiveContext::Queues::Code", 0)
    • After several seconds, verify that only 2 refs were processed by:

      • Check the remaining items in the queue (run ActiveContext::Queues.all_queued_items in the rails console)
      • Check the logs that BulkProcessWorker only ran once (run tail -f log/active_context.log in the Gitlab directory)

Optional: test queue values in simulated SaaS instance

  1. Set environment variable GITLAB_SIMULATE_SAAS=1

  2. Test queue values:

    Ai::ActiveContext::Queues::Code.limit_throughput?
    => false
    
    Ai::ActiveContext::Queues::Code.number_of_shards
    => 5 # the configured queue_shard_count
    
    Ai::ActiveContext::Queues::Code.shard_limit
    => 10000 # the configured queue_shard_limit

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #595434 (closed)

Edited by Pam Artiaga

Merge request reports

Loading
Loading