Fail Duo flows immediately when preconditions are not met

What this MR does and why

Duo Agent Platform flows execute as a CI job tagged gitlab--duo. When a project has no runner eligible to pick up that job (or its namespace has run out of compute minutes, which also removes instance runners), the job stays pending for up to an hour before CI reports it as stuck, and the session fails without a reason a user can act on.

This MR adds a precondition check at flow creation time. Ai::DuoWorkflows::CreateWorkflowService now calls Ai::DuoWorkflow::ProjectReadiness (an existing class, not new code) before starting the session. If the project is not ready, the service drops the session through UpdateWorkflowStatusService with a real, specific summary and returns an error to the caller immediately, instead of waiting for the CI timeout. Because CreateWorkflowService is the one entry point every flow trigger goes through, this single check covers all of them.

Implements https://gitlab.com/gitlab-org/gitlab/-/work_items/618844, which has the full problem writeup and supporting evidence.

How it works

  • The check only runs for sessions backed by a CI job. It is skipped for client-executed sessions (IDE, CLI, and create-without-start, identified by execution_mode), for agentic chat (which runs server-side rather than as a gitlab--duo job), and when there is no project to check.
  • Runner availability is checked first. Compute-minute exhaustion only removes instance runners, so a project with a usable top-level group runner can still run the job. Only when no runner is available do we use the minutes check to choose the right message (out of compute minutes vs no eligible runner).
  • The failure note now carries the actual reason for the drop instead of the generic "dropped" text. This is an additive reason: argument on UpdateWorkflowStatusService; existing callers that do not pass it are unaffected.
  • Code Review Flow now posts a specific note at creation time instead of a generic timeout note 30 minutes later. Two new reason keys were added to FailureMessageResolver, each with its own message and error code: DCR4011 for no eligible runner, DCR4012 for out of compute minutes. Troubleshooting docs were added alongside them.
  • If the drop itself cannot run (for example the status transition or permission check is refused), we record it through Gitlab::ErrorTracking rather than letting the session sit silently in created.

Feature flag

The check is behind duo_workflow_precondition_check (type gitlab_com_derisk), disabled by default, scoped to the project actor. With the flag off, behavior is unchanged from today. Rollout is tracked at #628626 (closed).

Testing

  • Unit specs cover the check itself (refusal when there is no runner, refusal when out of compute minutes, runner availability checked before minutes (including: out of minutes but a usable runner still creates the session), behavior with the flag off, and skipping for client-executed sessions), the new reason threading through the failure note, and the two new resolver keys.
  • The flag runs on by default in specs. Downstream caller specs that create a real session make the project look runner-ready through a shared helper (stub_duo_runner_available), so removing the flag later needs no further test changes.
  • Verified end to end on a local GDK, using a project with no gitlab--duo runner. With the flag on, creating a session returns an error naming the reason, the session ends in the failed state with the matching summary, and the merge request receives a start note followed by a note reading "session N failed (no instance or top-level group runner with the gitlab--duo tag)". With the flag off, the session stays in the created state, as it does today.

Notes for reviewers

  • Code Review only ever gets one note on refusal. The drop goes through update_workflow_system_note, which already checks suppress_agent_session_note?. Code Review sets that flag, so our drop note is suppressed there and Code Review's own resolver note is the only one posted.
  • Session funnel counting is not affected: agent_platform_session_created fires before this check runs, and agent_platform_session_dropped fires when we drop the session, so both events are still recorded.
  • The user-facing error strings now match the canonical troubleshooting section added in !252381 (merged), which documents the eligible-runner rule (the gitlab--duo tag, a Docker-compatible executor, and an instance or top-level group runner) and notes that running out of compute minutes only affects hosted runners. We dropped the earlier guidance about registering a project runner, and the @-mention reply and the DCR4011/DCR4012 troubleshooting entries link to that section so the messages and docs stay consistent. Further UX or technical-writing polish can still come in a followup, and since the flag defaults to off none of this is user-visible yet.
  • The query cost of this check at flow-start volume has not been measured yet. That is tracked as a gate on the rollout issue before the flag is enabled broadly.

Testing Screenshots

1. Fix Pipeline button (and Session)

Screenshot 2026-09-11 at 9.40.12 AM.png

Screenshot 2026-09-11 at 9.40.56 AM.png

2. Auto-trigger on an MR pipeline activity feed

Screenshot 2026-09-11 at 10.38.07 AM.png

4. Duo Code Review

Screenshot 2026-09-11 at 9.47.40 AM.png

5. Mention in a thread

Screenshot 2026-09-11 at 10.31.52 AM.png

Edited by Peter Gumeson

Merge request reports

Loading
Loading