Clear BatchLoader executor after ActionCable worker tasks

What does this MR do and why?

GraphQL subscription updates run on ActionCable worker threads. Those threads skip the Rack middleware and the Sidekiq middleware that clear BatchLoader's per-thread executor, so a thread that served one task keeps the values it batch-loaded and serves them again on the next one. Any subscription field resolved through BatchLoader::GraphQL can go stale this way. I hit it on the MR Pipelines tab, where failedJobsCount and retryable on ciPipelineStatusUpdated payloads sometimes reported the count from an earlier transition. The subscribe path fills the executor as well: authorize_object_or_gid! loads the object through GitlabSchema.find_by_gid, which batch-loads.

A new :work callback clears the executor after every ActionCable worker task (channel subscribe, receive, and each subscription update). It has the same shape as the existing load balancing callback and the same effect as the Sidekiq middleware per job.

This transport carries every GraphQL subscription, so the clear is behind the clear_action_cable_loader flag (gitlab_com_derisk, rollout in [FF] `clear_action_cable_loader` -- Clears the ... (#628832)). It consolidates Fix: Clear BatchLoader state on Action Cable su... (!255170 - closed), which fixed the same bug for the note session bar, as agreed in !255170 (comment 3831385720).

Split out of Keep MR failed jobs widget in sync with job ret... (!253179) at the reviewer's request so it can get backend review on its own. MR Pipelines tab: subscribe to rendered pipelin... (#628001) relies on it too.

How to set up and validate locally

The stale value is only visible in the subscription payloads themselves. On master the failed jobs widget polls its own count, so the UI won't show it directly.

  1. Enable mr_pipelines_graphql and ci_stage_subscription, leave clear_action_cable_loader disabled, open an MR whose pipeline has three or more failed jobs, and keep the Pipelines tab open.
  2. In the browser devtools, open the websocket connection under Network and watch the mrPipelineStatusUpdated frames.
  3. Retry the failed jobs one at a time from the stage dropdown. After each retry compare failedJobsCount in the incoming frames with Ci::Pipeline.find(id).latest_builds.failed.count in a Rails console. Some frames repeat the count from an earlier update, because the ActionCable thread that served it still holds the batch-loaded value. It depends on thread reuse, so it doesn't happen on every retry.
  4. Run Feature.enable(:clear_action_cable_loader) and repeat. Every frame matches the database.

The spec asserts the deterministic part: a value batched inside one worker task isn't reused by the next task on the same thread with the flag on, and is reused with it off.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Sahil Sharma

Merge request reports

Loading
Loading