Semantic Search: rate limit embeddings generation on GitLab Self-Managed

What does this MR do and why?

In https://gitlab.com/gitlab-org/modelops/applied-ml/code-suggestions/ai-assist/-/work_items/2085+, we decided that for Self-managed instances using GitLab-managed models, the indexing throughput should be limited to prevent abuse. This was implemented in !230480 (merged), but we had to override the Ai::ActiveContext::Code.limit_throughput? value to always return false since we could not provide the alternative option of Self-hosted models at that time.

In this MR:

  • update Ai::ActiveContext::Code.limit_throughput? return value, with the following scenarios:
    • on SaaS - not limited
    • on Self-Managed using their own Self-hosted model - not limited
    • on Self-Managed using the GitLab-managed model - limited
  • apply the throughput limiting logic to both the Code and CodeBackfill queues
    • the Code queue checks the current_indexing_embedding_model configuration
    • the CodeBackfill queue checks the next_indexing_embedding_model configuration

Throughput limiting logic:

Note: this was already implemented in !230480 (merged)

  1. If throughput is limited, the number_of_shards for a queue will always be 1
  2. If throughput is limited, the shard_limit (ie, number of references picked up for bulk processing) is limited according to a calculated max_shard_limit
  3. If throughput is limited, the BulkProcessWorker will not re-enqueue even if there are more items in the queue. It will instead wait to be re-run as a scheduled cron job (every minute).

UpdateQueueShardCountService and ReenqueueOrphanedRefs

  • if throughput is limited, shard count updates are not allowed since the collection_record.queue_shard_count value will not be respected anyway
  • ReenqueueOrphanedRefs task can still run, since these check the collection_class.queue.number_of_shards

References

Screenshots or screen recordings

N/A

How to set up and validate locally

For the return value of Ai::ActiveContext::Code.limit_throughput?, this should be covered in the unit tests.

For the throughput limiting logic (ie if limit_throughput? is true), see !230480 (merged).

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Related to #605813

Edited by Pam Artiaga

Merge request reports

Loading
Loading