Add rate limits for job retry and job play
What does this MR do and why?
Job retry and job play are expensive operations on GitLab.com. Across REST, GraphQL, and the web UI, they cost about 4.18 million server-seconds and 8.58 million Sidekiq jobs per day (24-hour sample, 2026-08-31 to 2026-09-01). Both are O(n) per call: one retry call reached 727 CI writes and 2,783 queries. One job received 2,878 play calls in 24 hours, 2,877 of which were rejected as unplayable. A small number of actors produce most of the rejected calls, so the rate-limit check runs before the authorization check, and rejected calls are counted too.
This MR adds four Labkit rate-limit rules in lib/gitlab/application_rate_limiter/labkit_adapter/supported_rate_limits.rb:
job_retry: per user and job, fixed limit of 3 per minute.job_retry_per_project: per user and project, configurable via the application settingjob_retry_limit_per_user_project, default 1250.job_play: per user and job, fixed limit of 3 per minute.job_play_per_project: per user and project, configurable via the application settingjob_play_limit_per_user_project, default 750.
Setting a per-project limit to 0 disables that limit. The per-job rules use a new ci_build characteristic; the services pass the job as ci_build: in the scope hash, so builds and bridges share the same buckets.
The default of 750 for play is the observed peak of 601 calls per minute plus headroom. The default of 1250 for retry applies the same ~25% headroom to the observed 1,019 calls per minute from a web UI actor. The issue asked to identify that actor first, so this default should be re-measured before the flag rolls out.
The checks live in Ci::RetryJobService#execute, Ci::PlayBuildService#execute and Ci::PlayBridgeService#execute (a played trigger job creates a downstream pipeline, so it shares the play buckets), before the retryable? and permission checks. The shared check, short-circuit and log entry live in Ci::JobRateLimitable; each service keeps its own literal feature flag check. The per-job check runs first and short-circuits the per-project check, so a caller already blocked on one job does not consume the project-wide budget. Throttled calls are logged with Gitlab::AppJsonLogger.
Exempted with rate_limit: false, because they are system-started:
Ci::Buildauto-retry after a failure.Deployments::AutoRollbackService(EE).- The retry fallback inside
Ci::PlayBuildService, so a play is not also counted as a retry. - Environment stop actions (
Environment#run_stop_action!, reached from auto-stop, merge request close and branch deletion as well as a manual stop), so a throttledon_stopjob cannot leave an environment stopped without running it.Ci::Build#playpassesrate_limit:through for this.
Two other callers of Ci::Build#play keep counting but no longer swallow a rejected play: Ci::PlayManualStageService logs it like an access error, and the ChatOps /deploy command answers with the reason instead of failing on a missing job.
Ci::RetryPipelineService uses clone!, not execute, and is not affected.
How it surfaces:
- REST:
POST /projects/:id/jobs/:job_id/retryandPOST /projects/:id/jobs/:job_id/playreturn429 Too Many Requestswith aRetry-Afterheader taken from the rule period. - GraphQL:
jobRetryandjobPlayreturn the message inerrors. - Web UI:
Projects::JobsController#retryand#playreturn 429 JSON for the pipeline graph action button, or a flash alert plus a redirect back to the job for the HTML flow.
Note that the controller and the REST endpoint reject a play of a job that is no longer playable? before the service runs, so a repeated play of the same job is stopped there with 400/422 rather than counted. The per-job play limit therefore mostly applies through GraphQL and the play fallback; the per-project limit covers the rest.
Both checks are behind gitlab_com_derisk feature flags, default off, milestone 19.5: rate_limit_job_retry (rollout #630118) and rate_limit_job_play (rollout #630119). Two flags so that play, which has firmer numbers, can roll out first.
Two admin settings are added under Admin > Settings > CI/CD > Continuous Integration and Deployment: "Maximum job retries per project" and "Maximum manual job runs per project". The Settings API and OpenAPI document are updated. Docs: doc/administration/cicd/limits.md (new section), doc/api/jobs.md (429), doc/api/settings.md, doc/user/gitlab_com/_index.md.
The same shape is used in !255403 (merged) (pipeline cancel) and !254367 (merged) (pipeline retry). Adjacent additions in supported_rate_limits.rb and labkit_adapter.rb mean a small merge conflict with those MRs is expected.
References
- Issue: https://gitlab.com/gitlab-org/gitlab/-/work_items/627273
- Rate-limit endpoint audit: https://gitlab.com/gitlab-org/gitlab/-/issues/605323
- Rollout issues: #630118, #630119
Screenshots or screen recordings
The only UI change is two new number fields on the admin CI/CD settings page.
| Before | After |
|---|---|
How to set up and validate locally
- In a Rails console, enable a flag:
Feature.enable(:rate_limit_job_play)(orrate_limit_job_retry). - Lower the limit:
ApplicationSetting.current.update!(job_play_limit_per_user_project: 1), or in Admin > Settings > CI/CD. - Run two manual jobs in the same project through the UI or
POST /api/v4/projects/:id/jobs/:job_id/play. - The second call is rejected:
429from the API, or a flash alert in the UI.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.