Productionise DAP suspendable workloads behind a feature flag
What does this MR do and why?
Productionises DAP (Duo Agent Platform) suspendable workloads behind the suspendable_environment_for_dap feature flag.
A DAP workload can now ask CI to suspend its environment when the job finishes, and a later workflow run can resume into that same environment instead of provisioning a fresh one.
The Ci-side half of that round trip — the runtime_environment_key param on PUT /api/v4/jobs/:id, the services that persist it, the read endpoint, and the environment_key → runtime_environment_key rename — landed on master in !239583 (merged) on 2026-09-08 and has been dropped from this MR. What remains:
- DAP threads suspension through.
StartWorkflowServicesetssuspend_on_success/suspend_on_failurefromsuspendable_environment_for_dap, andResumeWorkflowServicesupplies the previous run'sruntime_environment_keyso the new workload resumes into it. - The EE job router accepts the key too. !239583 (merged) added the param to
lib/api/ci/runner.rbonly;ee/lib/ee/api/internal/ci/job_router.rbhas its own copy ofPUT /jobs/:idand dropped it on the floor, so a job routed through the job router never round-tripped its key. ci_pending_builds.runner_machine_idgets its index and loose foreign key. The column and the queue filter that reads it are already onmaster(20260727185108,BuildQueueService#builds_for_runner_managerbehindci_pending_builds_runner_machine_id_filter); the partial index and theci_runner_machinesLFK entry are not.- The runtime environment key read moves to the runner API.
GET /projects/:id/jobs/:job_id/runtime_environment_key— added tolib/api/ci/jobs.rbby !239583 (merged) — is deleted and re-added asGET /api/v4/jobs/:id/runtime_environment_keyinlib/api/ci/runner.rb, authenticated byauthenticate_job!and nothing else, so only the CI job token belonging to that job is accepted. See notes for reviewers.
Closes #612876
Scope changed twice
This branch was originally a 4-commit stack that also added the ci_pending_builds.runner_machine_id column and the runner-machine queue filter; both landed on master independently and were removed here. The rebase onto post-!239583 master then removed the rest of the Ci-side work as well. Commit authorship is preserved: @ashvins for the index/LFK and the job-router param, @ssuman for the DAP side.
Runner requirements
Three things must line up on the runner side before a resume works. The Rails feature flag being on is not sufficient on its own, and none of the three surfaces an error when it is missing — the job suspends nothing, the next run starts cold, and nothing is logged as a failure.
1. Runner 19.4.0 or newer
The wire field was renamed environment_key -> runtime_environment_key in both directions by
gitlab-runner!6781 (merged)
(merged 2026-08-20, milestone 19.4, labelled missed:19.3):
| Direction | Runner source |
|---|---|
Server -> runner, in suspend_options |
common/spec/spec.go:412 |
Runner -> server, on PUT /api/v4/jobs/:id |
common/network.go:311 |
A runner on 19.3.x or older still reads and sends environment_key, so it receives nil and
never round-trips the key. v19.4.0 is now released and carries the rename together with the
docker executor's Suspend (executors/docker/docker_command.go:242) and resume path
(executors/docker/docker.go:1521), so a stock runner is sufficient.
2. FF_SUSPENDABLE_ENVIRONMENTS = true on the runner
[runners.feature_flags]
FF_SUSPENDABLE_ENVIRONMENTS = trueIt defaults to false (helpers/featureflags/flags.go), and nothing injects it as a job
variable — grepping app/, lib/ and ee/ for FF_SUSPENDABLE_ENVIRONMENTS returns nothing,
here or on master. Without it Build.WillSuspend() is false (common/build.go:288-293), so the
runner never calls suspendEnvironment, never reports a key, no Ci::RuntimeEnvironment row is
written, and the next run provisions a fresh environment. Reproduced on a GDK rig on 2026-09-17:
with the flag absent, Ci::JobRuntimeEnvironment rows still carried suspend_on_success: true
while every PUT /api/v4/jobs/:id arrived without runtime_environment_key, and the
ci_runtime_environments table stayed empty.
3. id set in the runner's config.toml
The key the runner builds is RuntimeEnvironmentKey{RunnerID: b.Runner.ID, SystemID: ...}
(common/build.go:313-318). With id unset or 0 the runner warns on every config load that
"suspend/resume jobs routed to this runner will fail when resumed". gitlab-runner register
writes it, but a hand-edited or regenerated config.toml loses it silently.
Feature flags
| Flag | Scope | Default |
|---|---|---|
suspendable_environment_for_dap |
DAP workload suspension (new in this MR) | off |
ci_suspendable_environment_runner_routing |
persisting the reported key, the read endpoint, runner-machine routing (already on master) | off |
How to set up and validate locally
- Enable the flags:
Feature.enable(:ci_suspendable_environment_runner_routing) Feature.enable(:suspendable_environment_for_dap) - Start a DAP workflow via the API and confirm the resulting CI job has a
Ci::JobRuntimeEnvironmentrow withsuspend_on_success: true. - When the job completes, the runner reports a
runtime_environment_key; verify aCi::RuntimeEnvironmentis created and linked (Ci::RuntimeEnvironments::RecordSuccessfulSuspensionService, onmaster). - Resume the workflow and confirm the new workload is dispatched with the same key, and that the runner skips
get_sources.
Verified end-to-end on a local GDK with a gitlab/gitlab-runner:bleeding docker runner: cold setup ~71s, resumed setup ~23s on the identical script.
Screenshots or screen recordings
N/A — backend-only change, no UI changes.
| Before | After |
|---|---|
| N/A | N/A |
Database queries
This MR adds no new queries. The writes that persist a reported key are on master
(Ci::RuntimeEnvironments::Record{Successful,Failed}SuspensionService), and
ResumeWorkflowService only traverses pipeline → builds → job_runtime_environment →
runtime_environment by primary key.
The schema change is one partial index:
CREATE INDEX index_ci_pending_builds_on_runner_machine_id
ON ci_pending_builds USING btree (runner_machine_id)
WHERE (runner_machine_id IS NOT NULL);It serves both the queue filter already on master and the loose-foreign-key cleaner, which
looks up ci_pending_builds rows by a specific runner_machine_id in order to nullify them.
spec/db/schema_spec.rb also requires it: it folds loose foreign keys into all_foreign_keys
and asserts every one is indexed, accepting a partial index only when its sole condition is the
column's presence, which is the shape of this one. Migration is byte-identical to the one
already covered by db:gitlabcom-database-testing.
Notes for reviewers
- The read endpoint is now job-token-only and lives in the runner API (@fabiopitino's two comments).
GET /projects/:id/jobs/:job_id/runtime_environment_keycame in from !239583 (merged) declaringpermissions: :update_jobwhile checking the TODO-deprecatedauthorize!(:update_build, build). An earlier revision of this MR swapped that for:read_job+authorize_read_builds!to match the siblingGETs — which widened the endpoint from developer to reporter, and that was the wrong direction for an executor-internal resume handle. It is now deleted fromlib/api/ci/jobs.rband re-added inlib/api/ci/runner.rbunderauthenticate_job!: only the owning job's token is accepted, becausejob_from_token'scurrent_job == found_jobcheck 403s another job's token and logsjob_token_mismatch. The route declares nopermissions:/boundary_type:, so it leaves the granular-token surface entirely — the generatedfine_grained_access_tokens_rest.mdnow lists it under CI job token — and it picks up the:runner_jobs_apirate limit the original lacked. Feature category stays:runner_core; the flag actor becomesjob.project, checked after authentication since runner routes have nouser_project. - Two consequences worth stating plainly, because they argue against the endpoint existing at all rather than against the move.
- Nothing consumes it. gitlab-runner only writes the key, on
PUT /api/v4/jobs/:id(common/network.go), and reads it back out ofsuspend_optionsin the job spec (common/spec/spec.go); it has no client method that GETs it.ResumeWorkflowServicereads the key in-process offjob_runtime_environment. - A job cannot read back the key produced by its own suspension.
authenticate_job!requiresEXECUTING_STATUSES(Ci::AuthJobFinder#validate_executing_job!), but that key is only written at job completion byCi::RuntimeEnvironments::RecordSuccessfulSuspensionServiceand its failed counterpart, so the suspending job itself always gets a 403. Pinned as a spec example (when the job has already finished→ 403). A resumed job is the reachable case:Gitlab::Ci::Pipeline::Chain::Create#bulk_insert_job_runtime_environments!links the new job'sCi::JobRuntimeEnvironmentto the existingCi::RuntimeEnvironment(looked up by key) at pipeline creation, so while that job runs the route returns 200 with the key of the environment it resumed into — also pinned as a spec example. The endpoint is therefore usable by a runner-side consumer on resume; it simply has none yet.
- Nothing consumes it. gitlab-runner only writes the key, on
- The old 6-example block in
spec/requests/api/ci/jobs_spec.rbis replaced byspec/requests/api/ci/runner/jobs_runtime_environment_key_get_spec.rb(14 examples): own token, token in theJOB-TOKENheader rather than the query string, another job's token (403 + mismatch log), no token, invalid token, erased job, missing job, finished job, flag disabled, rate limiting, application context. openapi_v{2,3}.yamlcarry only the moved path. A fullrake gitlab:openapi:v2:generateproduces ~3800 lines of unrelated pre-existing drift (/glql/schema,/iam/userinfo,/pipelines,/user/runners, response-description churn) andgitlab:openapi:v3:check_docsis run bystatic-analysis(scripts/static-analysis:69) while v2 is not, so I moved the single path block by hand and re-validated both files parse. That hand-move leftAPIEntitiesCiRuntimeEnvironmentKeyin a slot the generator does not emit, which failedstatic-analysis 2/2in pipeline 2858129292; the definition is now ordered as the generator emits it andgitlab:openapi:v3:check_docspasses locally.openapi_v2.yamlis unchecked by CI and its definition ordering may still differ from a full regeneration.config/routing/gitlab_routes.jsonis enforced bycells-routes:up-to-dateand was regenerated withgitlab:cells:routes:generate.cells-routes:router-in-syncis failing andpipeline:skip-router-synchas been applied; this is the justification. Unlike the blockingcells-routes:up-to-datejob above,router-in-syncisallow_failure: truein.gitlab/ci/cells.gitlab-ci.ymland, perdoc/development/cells/http_router.md, does not block merge requests today. It diffsconfig/routing/gitlab_routes.jsonbyte-for-byte against the HTTP Router's committed snapshot attest/routes/gitlab_routes.jsoningitlab-org/cells/http-router, and fails here because this MR moves a route — removing/api/:version/projects/:id/jobs/:job_id/runtime_environment_keyand adding/api/:version/jobs/:id/runtime_environment_key— which the router's snapshot doesn't reflect yet. Refreshing it needs a paired MR ingitlab-org/cells/http-routerrunningnpm run download-gitlab-routes, merging first before this pipeline re-runs clean; that cross-repo round trip is deferred rather than blocking this MR. The move is a straight substitution under the same/api/prefix, so it should match the router's existing rules and need only the refreshed snapshot, not a routing-rule change.config/authz/permissions/job/update.yml(update_job, added by !239583 (merged) for this endpoint) and itsassignable_permissionsentry are now referenced by no route. Left in place deliberately rather than reverted:update_jobshipped in 19.4, so removing an already-assignable granular permission would be breaking, andgitlab:permissions:validatepasses with it unreferenced. Runner core's call to deprecate.- The
Ai→Ciboundary leak @fabiopitino flagged inResumeWorkflowServiceis tracked in #628069.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.