Semantic Search: rate limit embeddings generation on GitLab Self-Managed
What does this MR do and why?
In https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085+, we decided that for Self-managed instances using GitLab-managed models, the indexing throughput should be limited to prevent abuse. This was implemented in !230480 (merged), but we had to override the Ai::ActiveContext::Code.limit_throughput? value to always return false since we could not provide the alternative option of Self-hosted models at that time.
In this MR:
- update
Ai::ActiveContext::Code.limit_throughput?return value, with the following scenarios:- on SaaS - not limited
- on Self-Managed using their own Self-hosted model - not limited
- on Self-Managed using the GitLab-managed model - limited
- apply the throughput limiting logic to both the
CodeandCodeBackfillqueues- the
Codequeue checks thecurrent_indexing_embedding_modelconfiguration - the
CodeBackfillqueue checks thenext_indexing_embedding_modelconfiguration
- the
Throughput limiting logic:
Note: this was already implemented in !230480 (merged)
- If throughput is limited, the
number_of_shardsfor a queue will always be1 - If throughput is limited, the
shard_limit(ie, number of references picked up for bulk processing) is limited according to a calculatedmax_shard_limit - If throughput is limited, the
BulkProcessWorkerwill not re-enqueue even if there are more items in the queue. It will instead wait to be re-run as a scheduled cron job (every minute).
UpdateQueueShardCountService and ReenqueueOrphanedRefs
- if throughput is limited, shard count updates are not allowed since the
collection_record.queue_shard_countvalue will not be respected anyway ReenqueueOrphanedRefstask can still run, since these check thecollection_class.queue.number_of_shards
References
- Issue: https://gitlab.com/gitlab-org/gitlab/-/work_items/605813+
- Original investigation: https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085+
Screenshots or screen recordings
N/A
How to set up and validate locally
For the return value of Ai::ActiveContext::Code.limit_throughput?, this should be covered in the unit tests.
For the throughput limiting logic (ie if limit_throughput? is true), see !230480 (merged).
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
Related to #605813