Draft: POC: Limit Code bulk processing throughput in SM instances

What does this MR do and why?

We are planning to make the rate limit rule stricter for the AIGW embeddings endpoint for requests coming from SM instances. This stricter rule could result in more rate limit errors, see https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085#note_3211312015. The increased rate limit errors need to be handled in Rails, by:

  1. Make the pipeline more robust so refs with rate limit errors can be retried indefinitely (#595433 (closed))

  2. Slow down the bulk processing queue for SM instances with Gitlab-operated code embeddings models - THIS MR

    Note: This is not the ultimate solution for rate limiting (that will be addressed on the infrastructure level on the AIGW endpoint), but once we do apply stricter rate limiting to the AIGW endpoints, we can pre-emptively reduce the number of 429 errors if we slow down bulk processing.

Proposed solutions

  • Limit queue throughputs - Set maximum values to the number_of_shards and shard_limit
  • Limit re-enqueuing
    • make sure the bulk processor does not re-enqueue
    • OR, increase the re-schedule interval

References

Screenshots or screen recordings

N/A

How to set up and validate locally

N/A

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #595434 (closed)

Edited by Pam Artiaga

Merge request reports

Loading
Loading