Gate release tags on the e2e matrix and integration suite

Description

v* tags shipped on unit tests and lint alone. The bash e2e matrix and the Go integration suite had no CI_COMMIT_TAG rule, and no publish job depended on them.

Stage split

Publish jobs were stage: build, the tests ran last, and needs: cannot point at a later stage, so gating needed a reorder:

test → build-binary → build-staging → e2e → publish → scan

Every Docker Hub image job, the swagger image and the CLI bucket uploads now sit in publish behind an explicit needs: on the full PG 10-18 matrix plus integration-test, shared through !reference [.e2e_gate, needs].

Images the tests pull

The e2e jobs pull ref-scoped images that tag pipelines never built, so the tag rule on its own would have failed every tag at docker pull:

  • build-image-{latest,rc}-server-dev now also push dblab-server:${CI_COMMIT_REF_SLUG}. These project-registry images are the candidate under test, so they stay in build-staging and are not gated on the tests that consume them.
  • build-image-tag-pg-upgrade builds pg-upgrade:<major>-${CI_COMMIT_REF_SLUG} for both tag shapes, which 6.clone_upgrade.sh and integration-test need.
  • build-binary-client{,-rc} had no artifacts: block and did their GCS upload inline. The e2e jobs install the CLI from engine/bin/cli/, so the build has to precede the tests and the upload has to follow them. Split into publish-binary-client{,-rc}.

Runner disk (#781)

Source data moves off /tmp, a half-of-RAM tmpfs on the dle-test runners, to /var/tmp/dle_test, overridable with DLE_TEST_DATA_DIR. This is the same path the manual runner-side symlink already points at, so nothing moves in practice.

_cleanup.sh had a line that looked like it cleared the source data but never matched anything, since both scripts override TMP_DATA_DIR to a subdirectory. The reuse is deliberate, so the dead line is gone and the intent is written down.

Stale source data

Second commit, from review feedback. The reused source cluster was never checked against the major about to run. The postgres entrypoint runs initdb only on an empty PGDATA, so a run killed part-way through it leaves a directory that can never start, and every later job for that major on the same runner inherits it. With tags now blocked on the matrix, one stuck directory holds up a release.

_source_data.sh compares PG_VERSION against POSTGRES_VERSION and drops the directory when it is missing or different. A valid directory is untouched, so the reuse it exists for is preserved. Existing stale directories on the runners are cleaned by the first job that touches them.

Both scripts also gain the readiness guard they were missing: the 300-second wait loop fell through silently, so a source database that never came up was reported by pgbench or by the pg_hba.conf edit that followed, far from its cause.

Toolchain

toolchain go1.26.8 pinned, CI images moved to match, golang.org/x/crypto v0.55.0 → v0.57.0.

govulncheck drops from 7 called vulnerabilities to 1: GO-2026-5932, the unmaintained golang.org/x/crypto/openpgp, reached only through the package init of go-github/v34 and with no fixed version available.

Orphaned e2e scripts

3.physical_walg.sh and 5.logical_rds.sh were invoked by nothing. Both need infrastructure CI does not have — a WAL-G archive that already holds a base backup, and a live RDS instance with IAM auth — and the dle-test shell executor cannot run GitLab services: at all, so the "wire it behind a minio service" option in the plan is not available as written.

Replaced by manual runbooks under docs/runbooks/, with engine/test/README.md recording why the numbering skips 3 and 5.

Not changed

publish-shared-release (npm) and ui_build_ce_image_release moved build → build-staging, which is the same position in the order they hold today. Gating UI releases on the engine e2e matrix is a policy change beyond this issue.

Related issue

#785

Depends on #781 for runner disk space. The repo-side half is in this MR; the runner-side symlink is already applied.

Examples

Pipeline graphs, simulated against the GitLab CI lint API with dry_run on a merged copy of the branch:

ref e2e publish
v4.2.0 9 × bash-test-main, integration-test, e2e-ce-ui-test 8 jobs, all after
v4.2.0-rc.3 same 11 5 jobs, all after
master 10 (no integration-test, as before) none — unchanged

The scripts derive their image tag from CI_COMMIT_REF_SLUG directly (TAG=${TAG:-${CI_COMMIT_REF_SLUG:-"master"}}), so a tag run pulls the artifact the pipeline just built rather than master.

Still to do before merge: push an rc tag and confirm the matrix passes on a tag. The simulation covers the graph shape, not that the jobs go green there.

Merge request reports

Loading
Loading