ci: parallelize pipeline jobs and stabilize the Go cache key
What
Three commits:
- Removes stage gating that serialized jobs with no dependency on each other, scopes
tests:integrationto the packages that hold integration tests, stabilizes the Go cache key, and marks the safe jobsinterruptible. - Groups jobs into
application-security-testingandchecksstages so the pipeline graph is readable, and drops the duplicate Gemnasium dependency scan. - Sets
workflow:auto_cancel:on_new_commit: interruptibleso superseded pipelines actually stop. - Moves the build gate to
build_test, the only expensive job in that chain.
No Go code changes.
Why
Measured across 19 recent pipelines (10 MR, 9 main) and 600 jobs. Queue time was zero throughout, so wall time was structural rather than capacity:
- Median wall time 13.8 min (MR), 13.6 min (main).
- The last job to finish in 18 of 19 pipelines was
dependency-scanning-profile-0, which runs for ~26s.
check_docs_update gated the two longest test jobs
.test declared no needs:, so lint, tests:unit and check_go_generated_code waited for the entire documentation stage. In 10 of 10 MR pipelines they started 2-3s after check_docs_update finished:
| pipeline | check_docs_update ends |
tests:unit starts |
|---|---|---|
| 2852344822 | 71s | 73s |
| 2852870829 | 167s | 168s |
| 2855009019 | 79s | 81s |
Mean dead wait before the longest test job even started: 123s. The SAST, dependency-scanning and secret-detection template jobs queued behind the same boundary. None of these consume another job's output.
The build stage waited on tests:integration
build_windows and snapcraft_build_edge also had no needs:. In pipeline 2855009019, tests:unit finished at t=234 and tests:integration at t=479; the build stage started at t=481. build_windows heads the pipeline's longest chain (windows_installer → build_test), so starting it 245s late pushed out the whole run.
The gate now sits at the end of that chain instead. build_windows (~60s) and windows_installer (~42s) run speculatively on medium runners, while build_test (~5 min on a 2xlarge) waits on tests:unit. Gating the head of the chain instead put tests:unit and the chain in series and cost ~2 min of wall time to protect two jobs cheap enough not to need it. snapcraft_build_edge stays gated: at ~178s it is never on the critical path, so gating it is free.
tests:integration had no cache and ran 370 packages to test 8
It extended only .test, not .go-cache; job traces go straight from Created fresh repository to go: downloading with no restore_cache section. It then ran -race -tags=integration -count=1 across all 370 packages, re-running the whole unit suite a second time, when only 8 packages contain a *_integration_test.go.
The Go cache key churned
Keyed on go.sum, so every dependency bump put each cache-using job on a key nobody had written yet, and this repo takes a bump every day or two:
| job | cache hit | fresh key |
|---|---|---|
tests:unit |
153s | 682s |
lint |
79s | 309s |
8 of 31 sampled lint runs landed on a fresh key. A per-job key is stable across bumps and makes the four hand-maintained cache prefixes redundant.
To be precise about what this does and does not buy, since a later pipeline on this branch measured it directly. It removes the key miss: the cache object now always exists, so the module cache is reused and unchanged packages still hit. It does not remove recompilation when a widely-imported dependency moves, because Go's build cache is content-addressed. Rebasing this branch picked up client-go v3.2.0 -> v3.4.0, which 462 files import; the cache was restored successfully and only 18 packages reported (cached), so tests:unit took 649s.
| no dependency change | after a dependency bump | |
|---|---|---|
| go.sum key (before) | 153s, when the key existed | 682-759s, key absent, modules re-downloaded |
| stable key (after) | 145-154s | 649s, key found, modules reused |
So on a dependency bump this is a modest improvement rather than an elimination. The large win applies to the pipelines where dependencies have not moved, which is most of them but never the one straight after a Renovate merge.
policy: pull-push is deliberately left alone. Routing writes to the default branch only looks tempting, but GitLab derives the -protected / -non_protected cache suffix from the triggering user's role, so every pipeline it classifies as non-protected (Developer-triggered MRs) would have gone permanently cold, since nothing else writes that slot. That penalty lands squarely on non-maintainer contributors, so it is not part of this change.
The test stage held twelve jobs of unrelated kinds
Stage order is now presentation only, since every job declares its own needs::
| stage | jobs |
|---|---|
application-security-testing |
gitlab-advanced-sast, secret_detection, and the policy-injected dependency-scanning-profile-0 |
checks |
lint_commit, check_go_version, check_embed, check_args, check_go_generated_code |
documentation |
unchanged |
test |
lint, lint:comments, tests:unit, tests:integration |
The security stage stays ahead of test because the policy-injected job's needs: is not ours to set, so its stage position is the only thing keeping it off the critical path. Its name is fixed by the policy for the same reason.
Superseded pipelines kept running
Marking jobs interruptible was not enough on its own. workflow:auto_cancel defaults to conservative, which refuses to cancel a superseded pipeline once any job with interruptible: false has started, and two such jobs start at t=1s that this project cannot mark: danger-review, from a component, and the policy-injected dependency-scanning job.
Observed on this branch: pushing a second commit left the first pipeline running its whole build stage, still going eleven minutes after it was superseded. on_new_commit: interruptible cancels the marked jobs and lets the rest finish.
on_job_failure is deliberately left at its default. Cancelling a pipeline on its first failure would discard the lint and documentation findings worth reading even when a test fails.
Two dependency scans ran per pipeline
Jobs/Dependency-Scanning.gitlab-ci.yml supplies the Gemnasium analyzer, deprecated in GitLab 17.9 and proposed for removal in 20.0, while an SBOM-based analyzer is already injected into every pipeline. The template include is dropped.
Effect
Measured on this MR's own pipelines, not modelled. Job timings from the fully warm run (pipeline 2855382183):
| job | baseline | now |
|---|---|---|
tests:unit |
303s median | 154s |
tests:integration |
315s median | 85s |
lint |
124s median | 83s |
check_docs_update |
gated the above | 69s, gates nothing |
dependency-scanning-profile-0 |
last job in the pipeline | finishes at 37s |
Resulting wall times:
| before | after | |
|---|---|---|
code-only MR (build stage skipped by changes:) |
7.1 min | 2.8 min |
MR touching .goreleaser.yml/.gitlab-ci.yml, main, merge train |
14.2 min | 7.7 min warm, 15.9 min straight after a dependency bump |
Both are measured on green pipelines. In the first case gitlab-advanced-sast (~170s) is now the critical path; in the second it is tests:unit into build_test, so the second figure tracks tests:unit and therefore cache state: 7.7 min when dependencies are unchanged, 15.9 min on the pipeline immediately after the client-go bump described above. The baseline 13.8 min median was itself measured across a mix of cache states. The full build-stage run (pipeline 2855401134, 7.7 min, passed):
2 145 tests:unit
2 73 build_windows: [amd64]
75 121 windows_installer: [amd64, x86_64]
147 459 build_testbuild_windows and windows_installer are done by 121s, so tests:unit at 145s is what build_test actually waits for. Shortening tests:unit is now the only remaining lever on this path.
Every job starts at t=1-2s. The new cache key is cold on its first run per job, then warms:
| pipeline | tests:unit |
tests:integration |
lint |
|---|---|---|---|
| 1st (cold, new key) | 759s | 211s | 211s |
| 2nd | 463s | 83s | 79s |
| 3rd | 154s | 85s | 83s |
Auto-cancel is also confirmed working: pipeline 2855382183, the first created with the new setting, cancelled itself when superseded, and the job it cancelled was build_test, the 2xlarge one. The two pipelines created before the setting existed had to be cancelled by hand, so the behaviour only applies from this change forward.
Plus roughly 8 runner-minutes per pipeline from the tests:integration scoping and the stable cache key, and another ~39s from dropping the duplicate dependency scan.
Notes for review
- Please sanity-check the Gemnasium removal. I confirmed empirically that
dependency-scanning-profile-0still runs after dropping the template, but I could not identify what injects it. It is not inglab ci config compileoutput, and thegitlab-orggroup's scan execution policies are allschedule-rule only, so none of them explains its presence on merge request pipelines. If that injection is less durable than it looks, or does not cover tag pipelines, dropping the in-repo scanner would be the wrong call. Happy to split this into its own MR if AppSec would rather review it separately. make integration-test-racenow defaults to the 8 integration packages.TEST_PKGS=on the command line still overrides it, andtest-raceis untouched. The list is derived withfindrather than hardcoded, so a new*_integration_test.gois picked up without editing the Makefile.interruptibleis opt-in and now covers every job that runs on a merge request, includingbuild_windows,windows_installer,build_testandsnapcraft_build_edge.release,snapcraft_release_*andreview-docs-*stay non-interruptible;build_testis marked on the job rather than on.release, whichreleasealso extends.code_navigation_golanghas noneeds:and so now waits on thechecksstage (~14s) in main pipelines. Left alone deliberately, since !3917 (merged) removes that job and an override here would break once it merges.- Worth a separate issue, not addressed here: merge request pipelines run by a maintainer write the same
-protectedGo cache slot thatmainand tag pipelines read, and that slot feedsbuild_windows→windows_installer→ the signed Windows installers that ship.