Feat(kubernetes): add suspend/resume support via single-PVC overlayfs
What does this MR do?
Enables Suspendable Environments feature on Kubernetes executor.
- Creates a PVC per job, mounts it as an overlayfs upper layer at
/persistand as the builds dir at/builds(viasubPath) - Runs an overlay entrypoint that
pivot_roots the container into the overlay before executing the job - After the job completes, deletes the pod but retains the PVC, emitting an environment key
On resume, the runner mounts the same PVC into the new pod — the prior container state (installed packages, files, builds dir) is visible immediately.
Why was this MR needed?
Part of the suspendable environments epic Resumable Jobs for CI and Agent Sessions (gitlab-org#21159) • Vishal Tak, Ashvin Sharma • 19.4. Kubernetes runner support.
Supported runtimes
- Direct host, AppArmor enforced (Ubuntu, GKE COS)
- RHEL
- Patched gVisor
- Kata Containers —
❌ not working
What's the best way to test this MR?
-
Create a Kubernetes cluster
-
Checkout this branch and configure Runner
Click to open
## Step 1 — Get a project and register a runner Run the following in a **single terminal session** and keep it open — `$ADMIN_TOKEN`, `$PROJECT_ID`, and `$RUNNER_ID` are needed in later steps. Use any existing project that has at least one commit on its default branch. The commands below find your first available project automatically: ```bash cd <PATH_TO_GITLAB> export ADMIN_TOKEN=$(bundle exec rails runner ' user = User.admins.first pat = user.personal_access_tokens.create!( name: "suspend-dev-pat", scopes: [:api], expires_at: 1.year.from_now ) puts pat.token ' 2>/dev/null | tail -1) # Pick any project with a default branch — or set PROJECT_ID manually export PROJECT_ID=$(curl -sf "http://gdk.test:3000/api/v4/projects?owned=true&per_page=10" \ -H "PRIVATE-TOKEN: $ADMIN_TOKEN" \ | python3 -c " import sys, json projects = [p for p in json.load(sys.stdin) if p.get('default_branch') and not p.get('empty_repo')] print(projects[0]['id'] if projects else '') ") echo "Using PROJECT_ID=$PROJECT_ID" RUNNER_INFO=$(curl -sf "http://gdk.test:3000/api/v4/user/runners" \ -X POST -H "PRIVATE-TOKEN: $ADMIN_TOKEN" -H "Content-Type: application/json" \ -d "{\"runner_type\":\"project_type\",\"project_id\":$PROJECT_ID, \"description\":\"k3d-suspend-runner\",\"tag_list\":\"k3d-suspend\",\"locked\":false}") export RUNNER_TOKEN=$(echo "$RUNNER_INFO" | python3 -c "import sys,json; print(json.load(sys.stdin)['token'])") export RUNNER_ID=$(echo "$RUNNER_INFO" | python3 -c "import sys,json; print(json.load(sys.stdin)['id'])") echo "RUNNER_ID=$RUNNER_ID" ``` --- ## Step 2 — Write config.toml ```bash RUNNER_WORK="$HOME/runner-suspend" mkdir -p "$RUNNER_WORK/config" cat > "$RUNNER_WORK/config/config.toml" << TOML concurrent = 1 log_level = "debug" log_format = "text" [[runners]] id = ${RUNNER_ID} name = "k3d-suspend-runner" url = "http://gdk.test:3000" token = "${RUNNER_TOKEN}" executor = "kubernetes" [runners.feature_flags] FF_SUSPENDABLE_ENVIRONMENTS = true [runners.kubernetes] namespace = "gitlab-runner" suspend_pvc_size = "1Gi" suspend_pvc_storage_class = "local-path" helper_image = "registry.gitlab.com/gitlab-org/gitlab-runner/gitlab-runner-helper:arm64-latest" [runners.kubernetes.build_container_security_context] privileged = true [runners.kubernetes.helper_container_security_context] privileged = true TOML ``` --- ## Step 3 — Start the runner ```bash RUNNER_BIN="<PATH_TO_RUNNER>/out/binaries/gitlab-runner-darwin-arm64" "$RUNNER_BIN" run \ --config "$RUNNER_WORK/config/config.toml" \ --working-directory "$RUNNER_WORK" \ > "$RUNNER_WORK/runner.log" 2>&1 & echo $! > "$RUNNER_WORK/runner.pid" sleep 3 grep "Checking for jobs" "$RUNNER_WORK/runner.log" | tail -2 # Expected: "Checking for jobs...no content" ``` -
Checkout
ashvins/send-env-key-with-job-completionon Rails -
Run this script. It will enable FF on rails, create 2 suspend jobs for one for success and failure attributes and exit.
Click to open
#!/usr/bin/env bash # scripts/suspend-e2e/rails-roundtrip.sh # # Orchestrator-perspective E2E round trip for suspend/resume, run once for # each suspend trigger (suspend_on_success, suspend_on_failure): # 1. Create a suspend build, print when it started # 2. Poll GET /jobs/:id for status (real API, no log-scraping) # 3. On the trigger firing, GET /jobs/:id/runtime_environment_key # 4. Create a resume build carrying that key, print when it started # 5. Poll GET /jobs/:id again, print when it completed # # There is currently no public REST endpoint for creating a pipeline with # suspend options set (checked: POST /projects/:id/pipeline does not accept # them) — only the internal Ci::CreatePipelineService does. create_pipeline() # below calls that service via `rails runner`, as a stand-in for the missing # orchestrator-facing endpoint. Everything else (status polling, key # retrieval) goes through the real REST API, matching what an orchestrator # holding only a PAT could actually do. # # No sleep, no log file scraping — status is read exclusively from # GET /jobs/:id. # # Standalone: everything needed is either a sane default below or derived # at runtime (PAT, project). Override any default by exporting the same-named # variable before running, e.g.: # GITLAB_URL=http://gdk.test:3000 RUNNER_TAG=my-tag ./rails-roundtrip.sh # # Usage: ./rails-roundtrip.sh set -euo pipefail export PATH="/opt/homebrew/bin:/opt/homebrew/sbin:${HOME}/.local/bin:${PATH}" SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" CI_YAML_SUCCESS_FILE="${SCRIPT_DIR}/.ci-roundtrip-success.yml" CI_YAML_FAILURE_FILE="${SCRIPT_DIR}/.ci-roundtrip-failure.yml" CREATE_PIPELINE_RB="${SCRIPT_DIR}/.create_pipeline.rb" # ── required (no defaults — export these before running) ──────────────────── GDK_GITLAB="${GDK_GITLAB:-}" GITLAB_URL="${GITLAB_URL:-}" GDK_PAT="${GDK_PAT:-}" PROJECT_ID="${PROJECT_ID:-}" RUNNER_TAG="${RUNNER_TAG:-suspend-e2e-test}" ts() { date '+%H:%M:%S'; } # Logs go to stderr: run_suspend_build/run_resume_build return their result # via stdout through $(...), and log lines must not end up captured in it. log() { echo "[$(ts)] $*" >&2; } die() { echo "[$(ts)] FATAL: $*" >&2 exit 1 } # ── setup ───────────────────────────────────────────────────────────────── check_dependencies() { for bin in curl jq; do command -v "$bin" >/dev/null || die "$bin not found in PATH" done } # The runner-side FF_SUSPENDABLE_ENVIRONMENTS gate is set in config.toml, but # suspend/resume also needs a Rails-side flag: without it, CreatePipelineService # silently drops the suspend_options this script passes in, and the round trip # fails with no obvious cause. Enable it here so this script is self-contained # and doesn't depend on having run preflight.sh first. enable_rails_feature_flag() { log "Enabling Rails feature flag ci_suspendable_environment_runner_routing..." local result result=$( cd "$GDK_GITLAB" export RAILS_ENV=development export DISABLE_SPRING=true mise exec -- bundle exec rails runner ' Feature.enable(:ci_suspendable_environment_runner_routing) puts Feature.enabled?(:ci_suspendable_environment_runner_routing) ' 2>&1 | tail -1 ) [[ "$result" == "true" ]] || die "failed to enable Rails FF ci_suspendable_environment_runner_routing. Got: ${result}" log "Rails FF ci_suspendable_environment_runner_routing: enabled" } # require_env: none of these have a safe default — assert the caller # exported them and fail immediately with clear instructions if not, # rather than guessing a path/host or auto-creating a PAT/project. require_env() { local missing=() [[ -n "$GDK_GITLAB" ]] || missing+=("GDK_GITLAB") [[ -n "$GITLAB_URL" ]] || missing+=("GITLAB_URL") [[ -n "$GDK_PAT" ]] || missing+=("GDK_PAT") [[ -n "$PROJECT_ID" ]] || missing+=("PROJECT_ID") [[ ${#missing[@]} -eq 0 ]] || die "required env var(s) not set: ${missing[*]}. Export them before running, e.g.: export GDK_GITLAB=\$HOME/workspace/gitlab-org/gdk/gitlab export GITLAB_URL=https://gdk.test:3000 export GDK_PAT=glpat-xxxxxxxxxxxxxxxxxxxx export PROJECT_ID=1" [[ -d "$GDK_GITLAB" ]] || die "GDK_GITLAB does not exist: ${GDK_GITLAB}" } # Writes both suspend-trigger variants of the CI YAML. Both detect which # phase they're in the same way (the builds-dir marker file left behind by # the suspend phase) and share an identical resume phase; they differ only # in whether the suspend phase exits 0 or 1, so suspend_on_success and # suspend_on_failure can each be exercised with their real, matching outcome. write_ci_yaml() { cat >"$CI_YAML_SUCCESS_FILE" <<'YAML' stages: - test suspend-resume-test: stage: test tags: - suspend-e2e-test script: - | if [ -f /builds/suspend-marker.txt ]; then echo "=== RESUME PHASE ===" cat /builds/suspend-marker.txt echo "RESUME_COMPLETE" else echo "=== SUSPEND PHASE (success) ===" echo "marker-$(date +%s)" > /builds/suspend-marker.txt echo "SUSPEND_PHASE_DONE" fi YAML cat >"$CI_YAML_FAILURE_FILE" <<'YAML' stages: - test suspend-resume-test: stage: test tags: - suspend-e2e-test script: - | if [ -f /builds/suspend-marker.txt ]; then echo "=== RESUME PHASE ===" cat /builds/suspend-marker.txt echo "RESUME_COMPLETE" else echo "=== SUSPEND PHASE (failure) ===" echo "marker-$(date +%s)" > /builds/suspend-marker.txt echo "SUSPEND_PHASE_DONE_WILL_FAIL" exit 1 fi YAML } # Writes the Ruby helper that create_pipeline() runs via `rails runner`. # Reads suspend options from environment variables rather than having them # interpolated into this string by bash — the runtime_environment_key # contains & and = characters that are miserable to shell-escape correctly. write_create_pipeline_script() { cat >"$CREATE_PIPELINE_RB" <<'RUBY' project = Project.find(ENV.fetch('E2E_PROJECT_ID').to_i) user = User.find_by_username('root') ci_yaml = File.read(ENV.fetch('E2E_CI_YAML_FILE')) suspend_options = if ENV['E2E_RUNTIME_ENVIRONMENT_KEY'] { runtime_environment_key: ENV['E2E_RUNTIME_ENVIRONMENT_KEY'] } else trigger = ENV.fetch('E2E_SUSPEND_TRIGGER', 'success') { suspend_on_success: trigger == 'success', suspend_on_failure: trigger == 'failure' } end result = Ci::CreatePipelineService.new(project, user, ref: 'main').execute( :web, content: ci_yaml, suspend_options: suspend_options ) pipeline = result.payload job = pipeline.builds.first puts "PIPELINE_ID=#{pipeline.id}" puts "PIPELINE_STATUS=#{pipeline.status}" puts "YAML_ERRORS=#{pipeline.yaml_errors}" puts "JOB_ID=#{job&.id}" RUBY } # ── GitLab API + Rails helpers ─────────────────────────────────────────── # create_pipeline [runtime_environment_key] [trigger_mode] -> prints JOB_ID to stdout # trigger_mode ("success"|"failure") selects the suspend-phase CI YAML and is # ignored when runtime_environment_key is set (a resume build's YAML is # identical either way). create_pipeline() { local env_key="${1:-}" local trigger_mode="${2:-success}" local yaml_file="$CI_YAML_SUCCESS_FILE" [[ -z "$env_key" && "$trigger_mode" == "failure" ]] && yaml_file="$CI_YAML_FAILURE_FILE" local output output=$( cd "$GDK_GITLAB" export RAILS_ENV=development export DISABLE_SPRING=true export E2E_PROJECT_ID="$PROJECT_ID" export E2E_CI_YAML_FILE="$yaml_file" export E2E_SUSPEND_TRIGGER="$trigger_mode" if [[ -n "$env_key" ]]; then export E2E_RUNTIME_ENVIRONMENT_KEY="$env_key" else unset E2E_RUNTIME_ENVIRONMENT_KEY fi mise exec -- bundle exec rails runner "$CREATE_PIPELINE_RB" ) local yaml_errors yaml_errors=$(echo "$output" | grep '^YAML_ERRORS=' | cut -d= -f2-) if [[ -n "$yaml_errors" && "$yaml_errors" != "[]" ]]; then die "pipeline YAML invalid: ${yaml_errors}" fi local job_id job_id=$(echo "$output" | grep '^JOB_ID=' | cut -d= -f2) [[ -n "$job_id" && "$job_id" != "" ]] || die "failed to create pipeline. Rails output: ${output}" echo "$job_id" } get_job_status() { curl -sk "${GITLAB_URL}/api/v4/projects/${PROJECT_ID}/jobs/$1" \ -H "PRIVATE-TOKEN: ${GDK_PAT}" | jq -r '.status' } # wait_for_terminal_status job_id -> prints final status to stdout. # No sleep: curl's own round-trip time is the only throttle between polls. wait_for_terminal_status() { local job_id="$1" status while true; do status=$(get_job_status "$job_id") case "$status" in success | failed | canceled) echo "$status" return ;; esac done } get_runtime_environment_key() { curl -sk "${GITLAB_URL}/api/v4/projects/${PROJECT_ID}/jobs/$1/runtime_environment_key" \ -H "PRIVATE-TOKEN: ${GDK_PAT}" | jq -r '.runtime_environment_key // empty' } # ── round trip phases ───────────────────────────────────────────────────── # run_suspend_build trigger_mode -> prints the retrieved runtime_environment_key to stdout # # A suspend_on_failure build is expected to *fail* — that's the trigger # firing correctly, not a script error. The job outcome is preserved across # suspend, so the expected terminal status tracks the trigger mode. run_suspend_build() { local trigger_mode="$1" local expected_status case "$trigger_mode" in success) expected_status="success" ;; failure) expected_status="failed" ;; *) die "unknown trigger mode: ${trigger_mode}" ;; esac log "Creating suspend build (trigger=${trigger_mode})..." local job_id status env_key job_id=$(create_pipeline "" "$trigger_mode") log "Build started: job_id=${job_id}" status=$(wait_for_terminal_status "$job_id") log "Build ${job_id} completed: status=${status}" [[ "$status" == "$expected_status" ]] || die "suspend build ended with status=${status}, expected ${expected_status} for trigger=${trigger_mode}" env_key=$(get_runtime_environment_key "$job_id") [[ -n "$env_key" ]] || die "no runtime_environment_key returned for job ${job_id}" log "Retrieved runtime_environment_key: ${env_key}" echo "$env_key" } # run_resume_build runtime_environment_key -> prints the resume job id to stdout run_resume_build() { local env_key="$1" log "Creating resume build..." local job_id status job_id=$(create_pipeline "$env_key") log "Build started: job_id=${job_id}" status=$(wait_for_terminal_status "$job_id") log "Build ${job_id} completed: status=${status}" [[ "$status" == "success" ]] || die "resume build did not succeed (status=${status})" echo "$job_id" } # run_round_trip trigger_mode -> runs one full suspend->resume cycle for that trigger run_round_trip() { local trigger_mode="$1" log "=== Trigger: ${trigger_mode} ===" local env_key resume_job_id env_key=$(run_suspend_build "$trigger_mode") resume_job_id=$(run_resume_build "$env_key") log "PASSED (${trigger_mode}): resume job=${resume_job_id}" } # ── main ────────────────────────────────────────────────────────────────── main() { check_dependencies require_env enable_rails_feature_flag write_ci_yaml write_create_pipeline_script run_round_trip success run_round_trip failure } main "$@" -
The script should be able to run without any issues.
What are the relevant issue numbers?
Resumable Jobs for CI and Agent Sessions (gitlab-org#21159) • Vishal Tak, Ashvin Sharma • 19.4
gitlab-runner: Add suspend/resume support to th... (#39464 - closed) • Ashvin Sharma • 19.4
Closes #39464 (closed)