ci: parallelize pipeline jobs and stabilize the Go cache key

What

Three commits:

  1. Removes stage gating that serialized jobs with no dependency on each other, scopes tests:integration to the packages that hold integration tests, stabilizes the Go cache key, and marks the safe jobs interruptible.
  2. Groups jobs into application-security-testing and checks stages so the pipeline graph is readable, and drops the duplicate Gemnasium dependency scan.
  3. Sets workflow:auto_cancel:on_new_commit: interruptible so superseded pipelines actually stop.
  4. Moves the build gate to build_test, the only expensive job in that chain.

No Go code changes.

Why

Measured across 19 recent pipelines (10 MR, 9 main) and 600 jobs. Queue time was zero throughout, so wall time was structural rather than capacity:

  • Median wall time 13.8 min (MR), 13.6 min (main).
  • The last job to finish in 18 of 19 pipelines was dependency-scanning-profile-0, which runs for ~26s.

check_docs_update gated the two longest test jobs

.test declared no needs:, so lint, tests:unit and check_go_generated_code waited for the entire documentation stage. In 10 of 10 MR pipelines they started 2-3s after check_docs_update finished:

pipeline check_docs_update ends tests:unit starts
2852344822 71s 73s
2852870829 167s 168s
2855009019 79s 81s

Mean dead wait before the longest test job even started: 123s. The SAST, dependency-scanning and secret-detection template jobs queued behind the same boundary. None of these consume another job's output.

The build stage waited on tests:integration

build_windows and snapcraft_build_edge also had no needs:. In pipeline 2855009019, tests:unit finished at t=234 and tests:integration at t=479; the build stage started at t=481. build_windows heads the pipeline's longest chain (windows_installer → build_test), so starting it 245s late pushed out the whole run.

The gate now sits at the end of that chain instead. build_windows (~60s) and windows_installer (~42s) run speculatively on medium runners, while build_test (~5 min on a 2xlarge) waits on tests:unit. Gating the head of the chain instead put tests:unit and the chain in series and cost ~2 min of wall time to protect two jobs cheap enough not to need it. snapcraft_build_edge stays gated: at ~178s it is never on the critical path, so gating it is free.

tests:integration had no cache and ran 370 packages to test 8

It extended only .test, not .go-cache; job traces go straight from Created fresh repository to go: downloading with no restore_cache section. It then ran -race -tags=integration -count=1 across all 370 packages, re-running the whole unit suite a second time, when only 8 packages contain a *_integration_test.go.

The Go cache key churned

Keyed on go.sum, so every dependency bump put each cache-using job on a key nobody had written yet, and this repo takes a bump every day or two:

job cache hit fresh key
tests:unit 153s 682s
lint 79s 309s

8 of 31 sampled lint runs landed on a fresh key. A per-job key is stable across bumps and makes the four hand-maintained cache prefixes redundant.

To be precise about what this does and does not buy, since a later pipeline on this branch measured it directly. It removes the key miss: the cache object now always exists, so the module cache is reused and unchanged packages still hit. It does not remove recompilation when a widely-imported dependency moves, because Go's build cache is content-addressed. Rebasing this branch picked up client-go v3.2.0 -> v3.4.0, which 462 files import; the cache was restored successfully and only 18 packages reported (cached), so tests:unit took 649s.

no dependency change after a dependency bump
go.sum key (before) 153s, when the key existed 682-759s, key absent, modules re-downloaded
stable key (after) 145-154s 649s, key found, modules reused

So on a dependency bump this is a modest improvement rather than an elimination. The large win applies to the pipelines where dependencies have not moved, which is most of them but never the one straight after a Renovate merge.

policy: pull-push is deliberately left alone. Routing writes to the default branch only looks tempting, but GitLab derives the -protected / -non_protected cache suffix from the triggering user's role, so every pipeline it classifies as non-protected (Developer-triggered MRs) would have gone permanently cold, since nothing else writes that slot. That penalty lands squarely on non-maintainer contributors, so it is not part of this change.

The test stage held twelve jobs of unrelated kinds

Stage order is now presentation only, since every job declares its own needs::

stage jobs
application-security-testing gitlab-advanced-sast, secret_detection, and the policy-injected dependency-scanning-profile-0
checks lint_commit, check_go_version, check_embed, check_args, check_go_generated_code
documentation unchanged
test lint, lint:comments, tests:unit, tests:integration

The security stage stays ahead of test because the policy-injected job's needs: is not ours to set, so its stage position is the only thing keeping it off the critical path. Its name is fixed by the policy for the same reason.

Superseded pipelines kept running

Marking jobs interruptible was not enough on its own. workflow:auto_cancel defaults to conservative, which refuses to cancel a superseded pipeline once any job with interruptible: false has started, and two such jobs start at t=1s that this project cannot mark: danger-review, from a component, and the policy-injected dependency-scanning job.

Observed on this branch: pushing a second commit left the first pipeline running its whole build stage, still going eleven minutes after it was superseded. on_new_commit: interruptible cancels the marked jobs and lets the rest finish.

on_job_failure is deliberately left at its default. Cancelling a pipeline on its first failure would discard the lint and documentation findings worth reading even when a test fails.

Two dependency scans ran per pipeline

Jobs/Dependency-Scanning.gitlab-ci.yml supplies the Gemnasium analyzer, deprecated in GitLab 17.9 and proposed for removal in 20.0, while an SBOM-based analyzer is already injected into every pipeline. The template include is dropped.

Effect

Measured on this MR's own pipelines, not modelled. Job timings from the fully warm run (pipeline 2855382183):

job baseline now
tests:unit 303s median 154s
tests:integration 315s median 85s
lint 124s median 83s
check_docs_update gated the above 69s, gates nothing
dependency-scanning-profile-0 last job in the pipeline finishes at 37s

Resulting wall times:

before after
code-only MR (build stage skipped by changes:) 7.1 min 2.8 min
MR touching .goreleaser.yml/.gitlab-ci.yml, main, merge train 14.2 min 7.7 min warm, 15.9 min straight after a dependency bump

Both are measured on green pipelines. In the first case gitlab-advanced-sast (~170s) is now the critical path; in the second it is tests:unit into build_test, so the second figure tracks tests:unit and therefore cache state: 7.7 min when dependencies are unchanged, 15.9 min on the pipeline immediately after the client-go bump described above. The baseline 13.8 min median was itself measured across a mix of cache states. The full build-stage run (pipeline 2855401134, 7.7 min, passed):

    2    145  tests:unit
    2     73  build_windows: [amd64]
   75    121  windows_installer: [amd64, x86_64]
  147    459  build_test

build_windows and windows_installer are done by 121s, so tests:unit at 145s is what build_test actually waits for. Shortening tests:unit is now the only remaining lever on this path.

Every job starts at t=1-2s. The new cache key is cold on its first run per job, then warms:

pipeline tests:unit tests:integration lint
1st (cold, new key) 759s 211s 211s
2nd 463s 83s 79s
3rd 154s 85s 83s

Auto-cancel is also confirmed working: pipeline 2855382183, the first created with the new setting, cancelled itself when superseded, and the job it cancelled was build_test, the 2xlarge one. The two pipelines created before the setting existed had to be cancelled by hand, so the behaviour only applies from this change forward.

Plus roughly 8 runner-minutes per pipeline from the tests:integration scoping and the stable cache key, and another ~39s from dropping the duplicate dependency scan.

Notes for review

  • Please sanity-check the Gemnasium removal. I confirmed empirically that dependency-scanning-profile-0 still runs after dropping the template, but I could not identify what injects it. It is not in glab ci config compile output, and the gitlab-org group's scan execution policies are all schedule-rule only, so none of them explains its presence on merge request pipelines. If that injection is less durable than it looks, or does not cover tag pipelines, dropping the in-repo scanner would be the wrong call. Happy to split this into its own MR if AppSec would rather review it separately.
  • make integration-test-race now defaults to the 8 integration packages. TEST_PKGS= on the command line still overrides it, and test-race is untouched. The list is derived with find rather than hardcoded, so a new *_integration_test.go is picked up without editing the Makefile.
  • interruptible is opt-in and now covers every job that runs on a merge request, including build_windows, windows_installer, build_test and snapcraft_build_edge. release, snapcraft_release_* and review-docs-* stay non-interruptible; build_test is marked on the job rather than on .release, which release also extends.
  • code_navigation_golang has no needs: and so now waits on the checks stage (~14s) in main pipelines. Left alone deliberately, since !3917 (merged) removes that job and an override here would break once it merges.
  • Worth a separate issue, not addressed here: merge request pipelines run by a maintainer write the same -protected Go cache slot that main and tag pipelines read, and that slot feeds build_windows → windows_installer → the signed Windows installers that ship.
Edited by Kai Armstrong

Merge request reports

Loading
Loading