Limit throughput of Code Embeddings Indexing pipeline in SM instances
What does this MR do and why?
We are planning to make the rate limit rule stricter for the AIGW embeddings endpoint for requests coming from SM instances. This stricter rule could result in more rate limit errors, see https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085#note_3211312015. The increased rate limit errors need to be handled in Rails, by:
-
Make the pipeline more robust so refs with rate limit errors can be retried indefinitely (#595433 (closed))
-
Slow down the bulk processing queue for SM instances with Gitlab-operated code embeddings models - THIS MR
Note: This is not the ultimate solution for rate limiting (that will be addressed on the infrastructure level on the AIGW endpoint), but once we do apply stricter rate limiting to the AIGW endpoints, we can pre-emptively reduce the number of
429errors if we slow down bulk processing.
Implemented solutions in this MR
- Limit queue throughputs - Set maximum values to the
number_of_shardsandshard_limit - Prevent re-enqueuing - make sure the bulk processor does not re-enqueue
References
- Issue: [Code Embeddings] Make the bulk processing slow... (#595434 - closed)
- Rate limiting investigations: https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085+
- PoC MR and throughput limiting discussions
Screenshots or screen recordings
N/A
How to set up and validate locally
Setup
-
Set the queue throughput configuration:
::Ai::ActiveContext::Collections::Code.collection_record.update_options!( queue_shard_count: 5, # anything higher than 1 queue_shard_limit: 10_000 # 1000 (default) or higher ) -
Set
ActiveContext::Config.re_enqueue_indexing_workers?to true. You can override this ingems/gitlab-active-context/lib/active_context/config.rb
Test in Self-Managed instance
-
Set environment variable
GITLAB_SIMULATE_SAAS=0 -
Test queue values:
Ai::ActiveContext::Queues::Code.limit_throughput? => true Ai::ActiveContext::Queues::Code.number_of_shards => 1 Ai::ActiveContext::Queues::Code.shard_limit => 600 -
Test the ref tracking:
# track refs ::Ai::ActiveContext::Collections::Code.track_refs!(routing: "1", hashes: [ "532997495b421186c8868e308b6fd1def6c9988beb76cb91508ea8738700aceb", "1311ee1cdd722cce9ac24a16c85212df40b69177b64fa353bbf4a205e6f25016", "9f3f66bd3e976a8b50d2cc94ff3bd1d0fd77ac5945f17ca6cf9e55e8925009f4", "56c38a4c7df12c221becfa62c410d19746982a9117abc6f8b732d384e100ee2c", "29d4bc8ed23583e6d481b18b62f6022f42b9f3df402ec0d2f6b7afb9302d2eeb", "6993a86dbb5ea3979cfbc10b6cd4b269ec4702ae8ce742324ceee581617367e6", "90ea0421f39ef710ba28f379151e9e87f6c1ceb915aa468b9b5eb732768bc4d4", "1229f89f0cc02ec87f16f09649840783786ca4d415d4948322c201b2012a418c", "c4087b44b8b5b3030da33a1c5a7490ed02b536202b349b15155308dfcae632a4", "8f38deab412aee170b4a81b36c6626b85c7d420cdc1e710e4358d52fb4e480e3" ]) # all refs are queued to 1 shard ActiveContext::Queues.all_queued_items => {"ai_activecontext_queues:{code}:0:zset"=> ["Ai::ActiveContext::References::Code|2|1|532997495b421186c8868e308b6fd1def6c9988beb76cb91508ea8738700aceb", "Ai::ActiveContext::References::Code|2|1|1311ee1cdd722cce9ac24a16c85212df40b69177b64fa353bbf4a205e6f25016", "Ai::ActiveContext::References::Code|2|1|9f3f66bd3e976a8b50d2cc94ff3bd1d0fd77ac5945f17ca6cf9e55e8925009f4", "Ai::ActiveContext::References::Code|2|1|56c38a4c7df12c221becfa62c410d19746982a9117abc6f8b732d384e100ee2c", "Ai::ActiveContext::References::Code|2|1|29d4bc8ed23583e6d481b18b62f6022f42b9f3df402ec0d2f6b7afb9302d2eeb", "Ai::ActiveContext::References::Code|2|1|6993a86dbb5ea3979cfbc10b6cd4b269ec4702ae8ce742324ceee581617367e6", "Ai::ActiveContext::References::Code|2|1|90ea0421f39ef710ba28f379151e9e87f6c1ceb915aa468b9b5eb732768bc4d4", "Ai::ActiveContext::References::Code|2|1|1229f89f0cc02ec87f16f09649840783786ca4d415d4948322c201b2012a418c", "Ai::ActiveContext::References::Code|2|1|c4087b44b8b5b3030da33a1c5a7490ed02b536202b349b15155308dfcae632a4", "Ai::ActiveContext::References::Code|2|1|8f38deab412aee170b4a81b36c6626b85c7d420cdc1e710e4358d52fb4e480e3"]} -
Test re-enqueue
-
in order to test locally, we need to set a small shard limit.
# this will be prioritized over the calculated shard limit since it's smaller ::Ai::ActiveContext::Collections::Code.collection_record.update_options!(queue_shard_limit: 2) -
Process the refs queued in Step 4:
::Ai::ActiveContext::BulkProcessWorker.new.perform("Ai::ActiveContext::Queues::Code", 0) -
After several seconds, verify that only 2 refs were processed by:
- Check the remaining items in the queue (run
ActiveContext::Queues.all_queued_itemsin the rails console) - Check the logs that
BulkProcessWorkeronly ran once (runtail -f log/active_context.login the Gitlab directory)
- Check the remaining items in the queue (run
-
Optional: test queue values in simulated SaaS instance
-
Set environment variable
GITLAB_SIMULATE_SAAS=1 -
Test queue values:
Ai::ActiveContext::Queues::Code.limit_throughput? => false Ai::ActiveContext::Queues::Code.number_of_shards => 5 # the configured queue_shard_count Ai::ActiveContext::Queues::Code.shard_limit => 10000 # the configured queue_shard_limit
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
Related to #595434 (closed)