CI: Remove the benchmark pipeline
This removes the benchmark CI stage, which has never produced a result and spent roughly nine hours of runner time a week not producing one. The perf-data branch that bench:store-results writes and detect_regressions.py reads does not exist, and bench:analyze was skipped or canceled in every scheduled run back to 2026-07-02. Repairing it would not make the measurements usable, for a reason that is a property of the runners rather than of the code.
Why not repair it
The jobs run on shared runners whose hardware varies between runs. The same job, bench:run:x86-64:sse, drew these two hosts two days apart:
| 2026-09-06 | 2026-09-08 | |
|---|---|---|
| Reported clock | 2450 MHz | 2696 MHz |
| L1 data cache | 32 KiB | 48 KiB |
| L2 cache | 512 KiB | 1 MiB |
Blocked GEMM, TRSM and GEMV throughput follows the cache hierarchy, so a Welch t-test comparing one run against 30 stored ones at the configured 5% threshold largely resolves which host the job drew. bench:regression-gate fails the pipeline whenever that test fires, so a version that ran to completion would mostly turn the weekly pipeline red on host noise, which is worse than no signal.
The immediate failure is separate and also unfixed by a timeout bump. On 2026-09-06 all three bench:run jobs hit the 3 h job_execution_timeout; a timeout discards the results artifact as well, so those runs left only a truncated trace.
What is removed, and what stays
Removed: the stage and its include, ci/benchmark.gitlab-ci.yml, build.benchmark.sh, run.benchmark.sh, detect_regressions.py, push_perf_data.sh, and benchmark_targets.txt.
Untouched: everything under benchmarks/ and unsupported/benchmarks/. Those remain the basis for local measurement and for a standalone suite, which is the direction of !2903 and #3112.
The weekly schedule also stays. Its other 91 jobs are the full test tier, the docs deploy and the weekly tag; only the 9 benchmark jobs go away. Nothing is left over on the GitLab side either: no project variable matches benchmark or performance, and the perf-data branch never existed to clean up.
.agents/benchmarking.md pointed at the two deleted scripts as the description of the "scheduled build and result format". That sentence now records the fact an agent actually needs, which is that no CI job builds or runs benchmarks, so the pipeline validates neither a benchmark's compilation nor a performance claim.
Replacement
#3146 proposes the CI-side successor: a per-merge-request harness that selects the benchmarks a diff can reach, builds the base and the head, and runs them interleaved on one machine in a single job. Interleaving is the point. A ratio measured against a co-resident control survives a host that a cross-run comparison cannot control for.
Validation
merged CI config -> POST /projects/15462818/ci/lint -> valid: true, 0 errors, 0 warnings
reuse lint, tracked files, before: 2038 / 2038 compliant
reuse lint, tracked files, after: 2032 / 2032 compliantThe REUSE count falls by exactly the six deleted files. Both lints ran against a clean git archive export rather than the working tree, since untracked local build directories otherwise dominate the output. A tree-wide search for the removed filenames returns nothing, and stage: benchmark appeared only in the deleted file.
Appendix: scheduled-run outcomes and the host-variance evidence
Every scheduled run of the stage:
2026-09-06 bench:run x3 failed job_execution_timeout (3 h each, no results artifact)
2026-08-30 bench:run failed script_failure, remaining bench jobs canceled
2026-08-02 bench:run x3 failed script_failure
2026-07-02 bench:run x4 failed script_failureGET /projects/15462818/repository/branches/perf-data returns 404.
Google Benchmark context blocks behind the table, same job name, two days apart:
job 16330045743 (2026-09-06, scheduled)
"host_name": "runner-jjqcuhrnc-project-15462818-concurrent-0",
"num_cpus": 32,
"mhz_per_cpu": 2450,
"size": 32768, <- L1d
"size": 524288, <- L2
job 16357140153 (2026-09-08, web)
"host_name": "runner-3sk86exer-project-15462818-concurrent-0",
"num_cpus": 32,
"mhz_per_cpu": 2696,
"size": 49152, <- L1d
"size": 1048576, <- L2The CPU models are not quoted because run.benchmark.sh captured the model name only into the combined results file, which the timeouts meant was never written.