Retry rate-limited ActiveContext batches until they succeed
Description
RateLimitError is temporary and self-resolving, but the retry chain treats it like a permanent bug. A rate limit that outlasts the chain sends the batch to the dead queue, where it waits for manual replay. This happened on GitLab.com on 2026-08-29: the AI Gateway rate-limited one batch of 977 code refs for more than 11 hours, and the batch dead-lettered.
This MR routes rate-limited batches back to RetryQueue from any queue. They retry every 5 minutes until the rate limit clears, and they never reach the dead queue. Other errors keep walking the chain as before.
Failures (unchanged):
Main queue ──fail──▶ RetryQueue (+5 min) ──fail──▶ SecondRetryQueue (+30 min)
──fail──▶ ThirdRetryQueue (+2 h) ──fail──▶ FourthRetryQueue (+8 h) ──fail──▶ DeadQueueRate limits (new):
Any queue ──rate limit──▶ RetryQueue (+5 min) ──rate limit──▶ RetryQueue (+5 min) ──▶ ... until it succeedsThe preprocessor returns rate-limited refs in their own result bucket, and every queue routes that bucket to RetryQueue. The caller declares which errors count as rate limits: Ai::ActiveContext::References::Code declares RateLimitError only.
This cannot repeat the retry loop that !251318 (merged) removed. That bug had two causes: retried items went back to a queue with no delay, and they did not count as failures, so the worker re-enqueued the shard every second. Now the items go to a delayed queue, and they count as failures in the worker result. The retry load is also bounded: a rate-limited batch fails on its first throttled request, and RetryQueue has one shard that the cron worker processes once per minute.
Related to #627333 (closed).