Add a rate limit to pipeline retry
What does this MR do and why?
Adds two rate limits to pipeline retry:
| Key | Scope | Limit |
|---|---|---|
pipeline_retry |
user + pipeline | fixed, 5/minute |
pipeline_retry_per_project |
user + project | pipeline_retry_limit_per_user_project setting, default 200/minute |
Why: retrying a pipeline clones every retryable build inside the request, so one call does work proportional to the pipeline's size, and the caller picks the pipeline. There was no dedicated limit, only the generic front-door throttle. Both buckets are keyed on the calling user.
The REST endpoint and the pipelineRetry GraphQL mutation are counted in Ci::RetryPipelineService#execute, which both reach. The web UI is counted in Projects::PipelinesController#retry instead, because that path answers the request and then does the work in Ci::RetryPipelineWorker; the worker opts out with rate_limit: false so each click is counted once. Counting in the controller keeps the decision atomic and synchronous, so a throttled click gets a 429 rather than a 204 it never acted on.
The limit is counted before the access check and reported after it, per https://gitlab.com/gitlab-org/gitlab/-/issues/627233: a rejected call still consumes budget, but the caller keeps the specific error rather than a throttle message.
No database migration, since the setting lives in the existing rate_limits JSONB column. Behind the rate_limit_pipeline_retry feature flag, off by default. REST returns 429 and GraphQL puts the message in errors. No changelog entry, since the flag is default-off.
Review feedback applied
Round 1: the check now sits in the controller for the UI path rather than peeking; blocked per-pipeline calls no longer drain the per-project bucket; a nil user is guarded; 429 is in the REST failure list, the OpenAPI spec and doc/api/pipelines.md; the admin help text no longer hardcodes the per-pipeline value; the docs say the limit does not cover automatic job retries.
Round 2: rejected calls count against the bucket again, matching the two acceptance criteria in https://gitlab.com/gitlab-org/gitlab/-/issues/627233, with the access error still returned first. The rate_limit opt-out is read once in the constructor rather than out of params, so no request parameter can switch a limit off. The 429 reason phrase is now Too Many Requests, matching the sibling cancel endpoint in the same generated file.
Worth flagging
The rationale quoted in https://gitlab.com/gitlab-org/gitlab/-/issues/627233 for counting before authorization does not hold for a service-level check. The REST endpoint calls authorize! :update_pipeline before pipeline.retry_failed, and the mutation calls authorized_find! before execute, so API-layer rejections never reach check_rate_limit in either ordering. What the ordering actually captures is the EE identity-verification and merge-train rejections raised inside check_access. The code now matches the acceptance criteria either way, but the issue's stated reason is worth correcting.
References
Both confidential:
- https://gitlab.com/gitlab-org/gitlab/-/issues/627233
- https://gitlab.com/gitlab-org/gitlab/-/issues/605323
Screenshots or screen recordings
There is a new "Maximum pipeline retry rate" field in Admin > Settings > CI/CD; screenshot still to be attached.
How to set up and validate locally
- Enable the flag:
Feature.enable(:rate_limit_pipeline_retry). - Set the limit:
ApplicationSetting.current.update!(pipeline_retry_limit_per_user_project: 1). - POST twice to
/api/v4/projects/:id/pipelines/:pipeline_id/retry. - The second request returns 429.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.