Draft: POC: Limit Code bulk processing throughput in SM instances
What does this MR do and why?
We are planning to make the rate limit rule stricter for the AIGW embeddings endpoint for requests coming from SM instances. This stricter rule could result in more rate limit errors, see https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085#note_3211312015. The increased rate limit errors need to be handled in Rails, by:
-
Make the pipeline more robust so refs with rate limit errors can be retried indefinitely (#595433 (closed))
-
Slow down the bulk processing queue for SM instances with Gitlab-operated code embeddings models - THIS MR
Note: This is not the ultimate solution for rate limiting (that will be addressed on the infrastructure level on the AIGW endpoint), but once we do apply stricter rate limiting to the AIGW endpoints, we can pre-emptively reduce the number of
429errors if we slow down bulk processing.
Proposed solutions
- Limit queue throughputs - Set maximum values to the
number_of_shardsandshard_limit - Limit re-enqueuing
- make sure the bulk processor does not re-enqueue
- OR, increase the re-schedule interval
References
- Issue: [Code Embeddings] Make the bulk processing slow... (#595434 - closed)
- Rate limiting investigations: https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085+
Screenshots or screen recordings
N/A
How to set up and validate locally
N/A
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
Related to #595434 (closed)