Draft: Cancel redundant pipelines from a redis cache
What does this MR do and why?
Ci::PipelineCreation::CancelRedundantPipelinesService finds pipelines to auto-cancel by querying p_ci_pipelines for a project and ref, filtered by status with a LIMIT, then queries p_ci_builds to check interruptibility. Both are scans or joins against large partitioned tables, on a hot path.
PostgreSQL 18's appendrel planner change (commit fbc0fe9a2e) collapses row estimates for that query shape. The planner drops the LIMIT early-exit and prices the plan to completion.
This MR is a proof of concept that replaces candidate discovery with a bounded Redis cache, so the hot path stops touching the partitioned tables.
Two things reviewers must not miss
1. One deliberate semantic change: conservative protection becomes monotonic.
Today the interruptibility check is status IN (running, success, failed). A non-interruptible build that ran and was later canceled stops protecting its pipeline. The new query asks started_at IS NOT NULL across all builds of the pipeline, including retried rows. Once a non-interruptible build has started, the pipeline stays protected for good.
This matches what the documentation promises, and it lets the cache hold the answer instead of re-querying on every pass. Please confirm this is acceptable.
2. There is no benchmark yet.
The performance claim is asserted structurally (no discovery query, no interruptibility query, one Redis call per take) and not measured. Nothing here compares wall-clock or plan cost against the old query.
Design
The new flow, when the flag is on:
Ci::InitialPipelineProcessWorkercalls a newCi::PipelineCreation::TrackRedundantPipelineServiceafterStartPipelineService. It offers the pipeline to the cache and enqueues the existingCi::CancelRedundantPipelinesWorker. Offering after start means a pipeline that never starts is never offered, and the pipeline is in the cache before the cancellation job looks for it.PipelineProcessWorkercalls a newCi::PipelineProcessing::ReconcileRedundantPipelineServiceafterCi::ProcessPipelineService. Pipeline processing already runs once per burst of job transitions, so this adds no new Sidekiq job per job transition. That is why no new event or subscriber was added.Ci::CancelRedundantPipelinesWorkerbranches on the flag:CancelClaimedPipelinesServicewhen on, the existingCancelRedundantPipelinesServicewhen off.Gitlab::Ci::Pipeline::Chain::CancelPendingPipelinesno longer enqueues when the flag is on. Otherwise it could run before the pipeline is in the cache.
New classes
All under 100 lines.
| Class | Responsibility |
|---|---|
Ci::PipelineIdentity |
Value object for the (id, partition_id) composite key, so callers can reference a pipeline without loading it |
Ci::RedundantPipelineCandidates |
Domain API over the cache, with nested Keys, Scripts, and Store collaborators. Store takes injected redis, keys, TTL, and max-taken |
Ci::RedundantPipelineEligibility |
Whether a pipeline is tracked, cancellable, or protected once started |
Ci::StartedNonInterruptibleBuilds |
Whether a pipeline has a non-interruptible build that ever started |
Ci::RedundantPipelineClaim |
Whether a pipeline taken from the cache is still one the newer pipeline supersedes |
Redis
Gitlab::Redis::Cache. Two keys per project and ref, sharing one Redis Cluster hash tag: a sorted set of candidates scored by pipeline id, and a set of blocked pipelines. The ref is SHA256-hashed, because it is user-controlled and unbounded in length. TTL is 24 hours.
Two Lua scripts run through Labkit::Redis::Script. One takes superseded pipelines and optionally offers the caller. The other blocks a pipeline. Taking removes members in the same call that returns them, so two pipelines racing to supersede the same older pipeline cancel it once. A single call takes at most 3000 pipelines.
Cache-loss policy
The cache is disposable. When a key expires or is evicted, we cancel nothing. We do not fall back to a query, because losing the cache must not put the database back on the hot path.
This makes cancellation best-effort. Some redundant pipelines may keep running. False negatives are preferred over cancelling current work.
Child pipelines
A parent can finish while its children still run, and cancelling a finished parent does nothing. A new Ci::Pipeline#family_cancelable? asks whether any pipeline in the family is still cancelable. Each family member is then cancelled individually with cascade_to_children: false, rather than cascading from the parent. Each child therefore uses its own auto-cancel mode.
Behaviour preserved
- Project auto-cancel setting
- The existing
disable_cancel_redundant_pipelines_serviceops kill switch - Exact project and ref scoping
- Same-SHA and newer-pipeline protection
- Child pipelines do not initiate cancellation
- Non-CI sources such as
webideremain excluded - All three
workflow:auto_cancel:on_new_commitmodes:none,conservative,interruptible - Per-child cancellation modes
- Multi-project downstream pipelines remain out of scope
Feature flag
| Name | ci_redundant_pipeline_redis_candidates |
| Type | wip |
| Default in production | Off |
| Default in test environment | On |
| Scope | Project |
Tests
Roughly 200 examples.
- With the flag on,
ActiveRecord::QueryRecorderproves there is no project and ref discovery query againstp_ci_pipelines, and noNOT EXISTSinterruptibility query. RedisCommands::Recorderproves that taking N candidates costs exactly one Redis call.- An integration spec drives the real
Ci::CreatePipelineServiceand the workers with:sidekiq_inline. - These existing specs re-run green:
cancel_redundant_pipelines_service,atomic_processing_service,pipeline_spec,build_spec,cancel_pipeline_service.
Not yet done
This is a proof of concept.
- No benchmark comparing the new path against the old query. The performance claim is structural, not measured.
- A pipeline removed from the cache after finishing is not re-offered when it is retried. Accepted best-effort miss.
- No metrics or observability.
References
- Discussion issue: #621543 — "Rethink how we detect redundant pipelines to remove partitioned-table scans"
- Related query optimization issue: #601090
- Earlier scalability epic: &12419 (closed)
- PostgreSQL commit that changed appendrel row estimates: https://git.postgresql.org/cgit/postgresql.git/commit/?id=fbc0fe9a2e
How to set up and validate locally
Expand for setup steps and the value matrix
The flag is off by default, so enable it for your test project first.
-
In
rails console, enable the flag for the project you will use:Feature.enable(:ci_redundant_pipeline_redis_candidates, Project.find(ID)) -
Create a project and add a
.gitlab-ci.ymlwith an interruptible long-running job:sleepy: interruptible: true script: - sleep 600 -
Push a commit to the default branch. Wait for the pipeline to start running.
-
Push a second commit to the same branch.
-
Observe that the first pipeline is canceled.
To exercise the cache-loss path, flush the Redis cache between step 3 and step 4:
Gitlab::Redis::Cache.with { |redis| redis.flushdb }Value matrix
Only the feature flag varies. Everything else is the two-commit scenario above.
| Feature flag | Preconditions | Expected result |
|---|---|---|
| On | An older running pipeline exists on the same ref | The older pipeline is canceled |
| On | No older pipeline on the ref | Nothing is canceled |
| Off | An older running pipeline exists on the same ref | The existing search path cancels it, behaviour unchanged |
| On | Redis cache flushed after the first pipeline started | Nothing is canceled, and no discovery query runs against p_ci_pipelines |
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist.
Reviewers, please pay particular attention to:
- The monotonic conservative-protection change described above. It changes behaviour for a non-interruptible build that started and was later canceled.
- The absence of a benchmark. If measurement is required before this can move past proof of concept, say so.
- The cache-loss policy. Cancelling nothing on cache loss is deliberate, and it makes cancellation best-effort.