Keep MR failed jobs widget in sync with job retries

What does this MR do and why?

Fixes Stage dropdown retry does not update failed job... (#606907)

On the MR Pipelines tab with mr_pipelines_graphql enabled, retrying a failed job from the stage dropdown doesn't reliably update the failed jobs badge and expanded list. I reconcile the row and any open detail with a read after the action, made the pipeline row the owner of the count, and removed the subscription bookkeeping that depended on a pipeline being alive.

Why the widget gets stuck

The widget had its own count query, which polled every 10 seconds while its cached pipeline was active. Once the pipeline failed, polling stopped. Retrying from the stage dropdown refreshed the pipeline row, but the widget's separate count query never learned that the pipeline was active again.

There is also a distinction between updating the count and updating the list. Suppose a pipeline has failed jobs A and B. Retrying A creates a new attempt, A2, and marks A as retried. The badge should become 1, the failed list should contain only B, and the stage dropdown should show A2 instead of A. Updating the pipeline's count doesn't remove A from an already-cached failed-jobs connection.

Subscriptions alone don't cover every retry either. If another job is still running in the same stage, both the stage and pipeline can remain running throughout the retry. The failed-job membership changes without requiring an aggregate status transition, so there may be no pipeline status event to deliver that change.

What changed

  • Retries reconcile from a read after the action. Every mini graph consumer runs the shared job_retry document. The MR-specific document and the provide/inject that routed it to the shared action button are gone, along with the provide the downstream dropdown needed to switch it off. After a job action the wrapper reads the row and any open detail, which covers the retries that no status payload describes. The success toast sits in the component that ran the mutation.
  • The row owns the count. The table passes it to the widget, replacing the widget's count query. The widget mounts only with failures. The badge now includes bridge and external failures, matching the list.
  • Details subscribe while open. The failed list and stage dropdown use declarative subscriptions. They fetch on opening, then use cache-first so subscription writes don't trigger extra requests.
  • Finished rows stay subscribed. I replaced alive/forced-alive tracking with displayed pipeline IDs, including up to three downstreams per parent. This lets the row observe external retries and bring the widget back after another failure.
  • Background reads fall back to visibility-aware polling. Rows poll every 60 seconds, an open failed-jobs list every 10, and an open stage dropdown every 60, all through setupQueryPollingByVisibility so a hidden tab stops reading. Other actions keep their post-action reads, and the single-row read uses the pipeline's actual project for fork MRs. The pipelines-specific reconnect recovery utility is no longer part of this MR: a shared version belongs in the Apollo client configuration, so I'll propose it separately.
  • Backend work is split. Executor cleanup is in Clear BatchLoader executor after ActionCable wo... (!254146 - merged), which this MR targets. I added job-definition and pipeline preloading here to avoid per-job queries. The request and browser specs for these cases, including same-status retries, stage subscriptions disabled, and the widget disappearing and reappearing, run in the pipeline for this head.

References

Screenshots or screen recordings

These recordings are from the earlier revision already attached to this MR. They illustrate the original behavior and intended interaction, but aren't verification of the current head. The after recording still needs refreshing against the current head.

Before Earlier fix recording

Legacy comparison: REST failed-pipelines recording.

How to set up and validate locally

  1. Enable mr_pipelines_graphql and open an MR with two failed jobs and another job still running in the same stage. Expand the failed-jobs widget.
  2. Retry one failure from the stage dropdown. Check that the count drops, only the other failure remains in the widget, and the dropdown shows the new attempt. The widget should remain expanded while its count stays positive.
  3. Repeat from the widget with the stage dropdown open. Check that both views update. Repeat with ci_stage_subscription disabled to exercise the post-action reads without stage delivery.
  4. Retry the last failure. The widget should disappear. Let the new attempt fail and check that the widget reappears with the new attempt's ID. Repeat with a retry initiated outside this browser.
  5. Retry from the dropdown while the widget is collapsed, then expand the widget. Check that it shows the current failures. Exercise whole-pipeline retry separately using the pipeline row's retry button.
  6. Repeat the widget retry scenarios with mr_pipelines_graphql disabled. Also check a fork MR with a source-project pipeline and a pipeline whose failed list exceeds 100 jobs.
  7. With websocket delivery unavailable, check that the row, an open failed-jobs list, and an open stage dropdown still refresh on their poll while the tab is visible, and that an expanded widget survives those reads. Switch to another tab and back to confirm polling stops and resumes.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Sahil Sharma

Merge request reports

Loading
Loading