Draft: Cancel redundant pipelines from a redis cache

What does this MR do and why?

Ci::PipelineCreation::CancelRedundantPipelinesService finds pipelines to auto-cancel by querying p_ci_pipelines for a project and ref, filtered by status with a LIMIT, then queries p_ci_builds to check interruptibility. Both are scans or joins against large partitioned tables, on a hot path.

PostgreSQL 18's appendrel planner change (commit fbc0fe9a2e) collapses row estimates for that query shape. The planner drops the LIMIT early-exit and prices the plan to completion.

This MR is a proof of concept that replaces candidate discovery with a bounded Redis cache, so the hot path stops touching the partitioned tables.

Two things reviewers must not miss

1. One deliberate semantic change: conservative protection becomes monotonic.

Today the interruptibility check is status IN (running, success, failed). A non-interruptible build that ran and was later canceled stops protecting its pipeline. The new query asks started_at IS NOT NULL across all builds of the pipeline, including retried rows. Once a non-interruptible build has started, the pipeline stays protected for good.

This matches what the documentation promises, and it lets the cache hold the answer instead of re-querying on every pass. Please confirm this is acceptable.

2. There is no benchmark yet.

The performance claim is asserted structurally (no discovery query, no interruptibility query, one Redis call per take) and not measured. Nothing here compares wall-clock or plan cost against the old query.

Design

The new flow, when the flag is on:

  • Ci::InitialPipelineProcessWorker calls a new Ci::PipelineCreation::TrackRedundantPipelineService after StartPipelineService. It offers the pipeline to the cache and enqueues the existing Ci::CancelRedundantPipelinesWorker. Offering after start means a pipeline that never starts is never offered, and the pipeline is in the cache before the cancellation job looks for it.
  • PipelineProcessWorker calls a new Ci::PipelineProcessing::ReconcileRedundantPipelineService after Ci::ProcessPipelineService. Pipeline processing already runs once per burst of job transitions, so this adds no new Sidekiq job per job transition. That is why no new event or subscriber was added.
  • Ci::CancelRedundantPipelinesWorker branches on the flag: CancelClaimedPipelinesService when on, the existing CancelRedundantPipelinesService when off.
  • Gitlab::Ci::Pipeline::Chain::CancelPendingPipelines no longer enqueues when the flag is on. Otherwise it could run before the pipeline is in the cache.

New classes

All under 100 lines.

Class Responsibility
Ci::PipelineIdentity Value object for the (id, partition_id) composite key, so callers can reference a pipeline without loading it
Ci::RedundantPipelineCandidates Domain API over the cache, with nested Keys, Scripts, and Store collaborators. Store takes injected redis, keys, TTL, and max-taken
Ci::RedundantPipelineEligibility Whether a pipeline is tracked, cancellable, or protected once started
Ci::StartedNonInterruptibleBuilds Whether a pipeline has a non-interruptible build that ever started
Ci::RedundantPipelineClaim Whether a pipeline taken from the cache is still one the newer pipeline supersedes

Redis

Gitlab::Redis::Cache. Two keys per project and ref, sharing one Redis Cluster hash tag: a sorted set of candidates scored by pipeline id, and a set of blocked pipelines. The ref is SHA256-hashed, because it is user-controlled and unbounded in length. TTL is 24 hours.

Two Lua scripts run through Labkit::Redis::Script. One takes superseded pipelines and optionally offers the caller. The other blocks a pipeline. Taking removes members in the same call that returns them, so two pipelines racing to supersede the same older pipeline cancel it once. A single call takes at most 3000 pipelines.

Cache-loss policy

The cache is disposable. When a key expires or is evicted, we cancel nothing. We do not fall back to a query, because losing the cache must not put the database back on the hot path.

This makes cancellation best-effort. Some redundant pipelines may keep running. False negatives are preferred over cancelling current work.

Child pipelines

A parent can finish while its children still run, and cancelling a finished parent does nothing. A new Ci::Pipeline#family_cancelable? asks whether any pipeline in the family is still cancelable. Each family member is then cancelled individually with cascade_to_children: false, rather than cascading from the parent. Each child therefore uses its own auto-cancel mode.

Behaviour preserved

  • Project auto-cancel setting
  • The existing disable_cancel_redundant_pipelines_service ops kill switch
  • Exact project and ref scoping
  • Same-SHA and newer-pipeline protection
  • Child pipelines do not initiate cancellation
  • Non-CI sources such as webide remain excluded
  • All three workflow:auto_cancel:on_new_commit modes: none, conservative, interruptible
  • Per-child cancellation modes
  • Multi-project downstream pipelines remain out of scope

Feature flag

Name ci_redundant_pipeline_redis_candidates
Type wip
Default in production Off
Default in test environment On
Scope Project

Tests

Roughly 200 examples.

  • With the flag on, ActiveRecord::QueryRecorder proves there is no project and ref discovery query against p_ci_pipelines, and no NOT EXISTS interruptibility query.
  • RedisCommands::Recorder proves that taking N candidates costs exactly one Redis call.
  • An integration spec drives the real Ci::CreatePipelineService and the workers with :sidekiq_inline.
  • These existing specs re-run green: cancel_redundant_pipelines_service, atomic_processing_service, pipeline_spec, build_spec, cancel_pipeline_service.

Not yet done

This is a proof of concept.

  • No benchmark comparing the new path against the old query. The performance claim is structural, not measured.
  • A pipeline removed from the cache after finishing is not re-offered when it is retried. Accepted best-effort miss.
  • No metrics or observability.

References

How to set up and validate locally

Expand for setup steps and the value matrix

The flag is off by default, so enable it for your test project first.

  1. In rails console, enable the flag for the project you will use:

    Feature.enable(:ci_redundant_pipeline_redis_candidates, Project.find(ID))
  2. Create a project and add a .gitlab-ci.yml with an interruptible long-running job:

    sleepy:
      interruptible: true
      script:
        - sleep 600
  3. Push a commit to the default branch. Wait for the pipeline to start running.

  4. Push a second commit to the same branch.

  5. Observe that the first pipeline is canceled.

To exercise the cache-loss path, flush the Redis cache between step 3 and step 4:

Gitlab::Redis::Cache.with { |redis| redis.flushdb }

Value matrix

Only the feature flag varies. Everything else is the two-commit scenario above.

Feature flag Preconditions Expected result
On An older running pipeline exists on the same ref The older pipeline is canceled
On No older pipeline on the ref Nothing is canceled
Off An older running pipeline exists on the same ref The existing search path cancels it, behaviour unchanged
On Redis cache flushed after the first pipeline started Nothing is canceled, and no discovery query runs against p_ci_pipelines

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist.

Reviewers, please pay particular attention to:

  • The monotonic conservative-protection change described above. It changes behaviour for a non-interruptible build that started and was later canceled.
  • The absence of a benchmark. If measurement is required before this can move past proof of concept, say so.
  • The cache-loss policy. Cancelling nothing on cache loss is deliberate, and it makes cancellation best-effort.

Merge request reports

Loading
Loading