Fix runner pause state overwritten by a stale list refetch
What does this MR do and why?
(Hopefully) definitive fix for https://gitlab.com/gitlab-org/quality/test-failure-issues/-/work_items/44287.
The runner list queries (group and admin runner pages) use the Apollo cache-and-network fetch policy with no nextFetchPolicy, so every mutation that updates a runner also re-runs the list query over the network.
When a user pauses a runner and quickly resumes it (or the reverse), the resume-triggered refetch is absorbed into the still-in-flight pause-triggered refetch by Apollo query deduplication (on by default). That in-flight response carries a stale state before the resume. Nothing refetches after that, so the runner keeps a "Paused" badge until the page reloads.
This caused a flaky test at spec/features/groups/runners/owner_manages_runners_spec.rb (and a tiny UI bug).
The fix adds context: { queryDeduplication: false } to both runner list smart queries.
It also adds documentation to teach folks how/when to use this feature.
Screenshots or screen recordings
No visual change in normal use. The bug being fixed is itself visual: a runner row that keeps a "Paused" badge after a successful resume.
How to set up and validate locally
This is hard to validate manually! It is revealed most by a flaky test, but to reproduce the bug manually in a browser, hold the list responses so the race window is easy to hit:
-
Visit a group's runners page (Build > Runners) with at least one active runner.
-
In the DevTools console, delay every runner list response by 2 seconds:
const originalFetch = window.fetch; window.fetch = async (...args) => { let operationName; try { operationName = JSON.parse(args[1]?.body || '{}').operationName; } catch { // not a GraphQL request } const response = await originalFetch(...args); if (operationName?.startsWith('getGroupRunners')) { await new Promise((resolve) => setTimeout(resolve, 2000)); } return response; }; -
Click Pause on a runner, then click Resume within 2 seconds.
-
Without this fix, the "Paused" badge comes back after ~2 seconds and stays until reload. With this fix, the badge clears and stays cleared.
References
- https://gitlab.com/gitlab-org/quality/test-failure-issues/-/work_items/44287 — the flaky test issue this MR fixes
- https://gitlab.com/gitlab-org/quality/test-failure-issues/-/work_items/43855 — previous incarnation of the same flake
- https://gitlab.com/gitlab-org/quality/test-failure-issues/-/work_items/43165 — earlier incarnation of the flake
- !249594 (merged) — merged earlier fix for a different mechanism (dropped clicks on loading buttons)
- !249618 (merged) — merged earlier fix: Capybara treats aria-disabled buttons as disabled
- !252055 (merged) — open MR that quarantines the spec; not needed if this fix holds
- !252236 (merged) — split-out counter change that should merge before this MR
- https://gitlab.com/gitlab-org/gitlab/-/jobs/16098083526 — failing master scheduled-pipeline job whose graphql_json.log confirmed the mechanism