Add retry chain for ActiveContext failure handling
Description
The ActiveContext dead queue accumulates a large number of items that need manual intervention. Most of the errors we have seen come from transient AI Gateway incidents that last 1 to 13 hours, and one batch-level error can send up to 1,000 items to the dead queue at once. Today an item gets a single retry after 5 minutes, which cannot cover incidents of that length. Investigation: #606544 (comment 3710630691).
We introduce a chain of retry queues, as described in Option C of the original proposal in #604055 (closed).
Before:
Main queue ──fail──▶ RetryQueue (+5 min) ──fail──▶ DeadQueueAfter:
Main queue ──fail──▶ RetryQueue (+5 min) ──fail──▶ SecondRetryQueue (+30 min)
──fail──▶ ThirdRetryQueue (+2 h) ──fail──▶ FourthRetryQueue (+8 h) ──fail──▶ DeadQueueEach stage gives one retry, so an item gets ~10.6 hours of coverage before it reaches the dead queue. Each queue class points to its failure_queue, so the attempt count lives in the queue topology and items need no extra state. Broken items still reach the dead queue and surface to a human. There are no new Sidekiq workers and no migrations: the queues are Redis sorted sets, processed by the existing cron worker.
Other relevant changes in this MR:
RateLimitErrorleaves the infinite-retry list. Retried batches did not count as failures, so the worker re-enqueued the shard every second and retried against an already rate-limited gateway. These batches now walk the chain like any other failure.- The dead queue replay API accepts the new queues as targets.
Related to #606544 (closed).