Add retry chain for ActiveContext failure handling

Description

The ActiveContext dead queue accumulates a large number of items that need manual intervention. Most of the errors we have seen come from transient AI Gateway incidents that last 1 to 13 hours, and one batch-level error can send up to 1,000 items to the dead queue at once. Today an item gets a single retry after 5 minutes, which cannot cover incidents of that length. Investigation: #606544 (comment 3710630691).

We introduce a chain of retry queues, as described in Option C of the original proposal in #604055 (closed).

Before:

Main queue ──fail──▶ RetryQueue (+5 min) ──fail──▶ DeadQueue

After:

Main queue ──fail──▶ RetryQueue (+5 min) ──fail──▶ SecondRetryQueue (+30 min)
    ──fail──▶ ThirdRetryQueue (+2 h) ──fail──▶ FourthRetryQueue (+8 h) ──fail──▶ DeadQueue

Each stage gives one retry, so an item gets ~10.6 hours of coverage before it reaches the dead queue. Each queue class points to its failure_queue, so the attempt count lives in the queue topology and items need no extra state. Broken items still reach the dead queue and surface to a human. There are no new Sidekiq workers and no migrations: the queues are Redis sorted sets, processed by the existing cron worker.

Other relevant changes in this MR:

  • RateLimitError leaves the infinite-retry list. Retried batches did not count as failures, so the worker re-enqueued the shard every second and retried against an already rate-limited gateway. These batches now walk the chain like any other failure.
  • The dead queue replay API accepts the new queues as targets.

Related to #606544 (closed).

Merge request reports

Loading
Loading