Fail Duo flows immediately when preconditions are not met
What this MR does and why
Duo Agent Platform flows execute as a CI job tagged gitlab--duo. When a project has no runner eligible to pick up that job (or its namespace has run out of compute minutes, which also removes instance runners), the job stays pending for up to an hour before CI reports it as stuck, and the session fails without a reason a user can act on.
This MR adds a precondition check at flow creation time. Ai::DuoWorkflows::CreateWorkflowService now calls Ai::DuoWorkflow::ProjectReadiness (an existing class, not new code) before starting the session. If the project is not ready, the service drops the session through UpdateWorkflowStatusService with a real, specific summary and returns an error to the caller immediately, instead of waiting for the CI timeout. Because CreateWorkflowService is the one entry point every flow trigger goes through, this single check covers all of them.
Implements https://gitlab.com/gitlab-org/gitlab/-/work_items/618844, which has the full problem writeup and supporting evidence.
How it works
- The check only runs for sessions backed by a CI job. It is skipped for client-executed sessions (IDE, CLI, and create-without-start, identified by
execution_mode), for agentic chat (which runs server-side rather than as agitlab--duojob), and when there is no project to check. - Runner availability is checked first. Compute-minute exhaustion only removes instance runners, so a project with a usable top-level group runner can still run the job. Only when no runner is available do we use the minutes check to choose the right message (out of compute minutes vs no eligible runner).
- The failure note now carries the actual reason for the drop instead of the generic "dropped" text. This is an additive
reason:argument onUpdateWorkflowStatusService; existing callers that do not pass it are unaffected. - Code Review Flow now posts a specific note at creation time instead of a generic timeout note 30 minutes later. Two new reason keys were added to
FailureMessageResolver, each with its own message and error code: DCR4011 for no eligible runner, DCR4012 for out of compute minutes. Troubleshooting docs were added alongside them. - If the drop itself cannot run (for example the status transition or permission check is refused), we record it through
Gitlab::ErrorTrackingrather than letting the session sit silently increated.
Feature flag
The check is behind duo_workflow_precondition_check (type gitlab_com_derisk), disabled by default, scoped to the project actor. With the flag off, behavior is unchanged from today. Rollout is tracked at #628626 (closed).
Testing
- Unit specs cover the check itself (refusal when there is no runner, refusal when out of compute minutes, runner availability checked before minutes (including: out of minutes but a usable runner still creates the session), behavior with the flag off, and skipping for client-executed sessions), the new
reasonthreading through the failure note, and the two new resolver keys. - The flag runs on by default in specs. Downstream caller specs that create a real session make the project look runner-ready through a shared helper (
stub_duo_runner_available), so removing the flag later needs no further test changes. - Verified end to end on a local GDK, using a project with no
gitlab--duorunner. With the flag on, creating a session returns an error naming the reason, the session ends in thefailedstate with the matching summary, and the merge request receives a start note followed by a note reading "session N failed (no instance or top-level group runner with thegitlab--duotag)". With the flag off, the session stays in thecreatedstate, as it does today.
Notes for reviewers
- Code Review only ever gets one note on refusal. The drop goes through
update_workflow_system_note, which already checkssuppress_agent_session_note?. Code Review sets that flag, so our drop note is suppressed there and Code Review's own resolver note is the only one posted. - Session funnel counting is not affected:
agent_platform_session_createdfires before this check runs, andagent_platform_session_droppedfires when we drop the session, so both events are still recorded. - The user-facing error strings now match the canonical troubleshooting section added in !252381 (merged), which documents the eligible-runner rule (the
gitlab--duotag, a Docker-compatible executor, and an instance or top-level group runner) and notes that running out of compute minutes only affects hosted runners. We dropped the earlier guidance about registering a project runner, and the @-mention reply and the DCR4011/DCR4012 troubleshooting entries link to that section so the messages and docs stay consistent. Further UX or technical-writing polish can still come in a followup, and since the flag defaults to off none of this is user-visible yet. - The query cost of this check at flow-start volume has not been measured yet. That is tracked as a gate on the rollout issue before the flag is enabled broadly.




