Add rate limits for job retry and job play

What does this MR do and why?

Job retry and job play are expensive operations on GitLab.com. Across REST, GraphQL, and the web UI, they cost about 4.18 million server-seconds and 8.58 million Sidekiq jobs per day (24-hour sample, 2026-08-31 to 2026-09-01). Both are O(n) per call: one retry call reached 727 CI writes and 2,783 queries. One job received 2,878 play calls in 24 hours, 2,877 of which were rejected as unplayable. A small number of actors produce most of the rejected calls, so the rate-limit check runs before the authorization check, and rejected calls are counted too.

This MR adds four Labkit rate-limit rules in lib/gitlab/application_rate_limiter/labkit_adapter/supported_rate_limits.rb:

  • job_retry: per user and job, fixed limit of 3 per minute.
  • job_retry_per_project: per user and project, configurable via the application setting job_retry_limit_per_user_project, default 1250.
  • job_play: per user and job, fixed limit of 3 per minute.
  • job_play_per_project: per user and project, configurable via the application setting job_play_limit_per_user_project, default 750.

Setting a per-project limit to 0 disables that limit. The per-job rules use a new ci_build characteristic; the services pass the job as ci_build: in the scope hash, so builds and bridges share the same buckets.

The default of 750 for play is the observed peak of 601 calls per minute plus headroom. The default of 1250 for retry applies the same ~25% headroom to the observed 1,019 calls per minute from a web UI actor. The issue asked to identify that actor first, so this default should be re-measured before the flag rolls out.

The checks live in Ci::RetryJobService#execute, Ci::PlayBuildService#execute and Ci::PlayBridgeService#execute (a played trigger job creates a downstream pipeline, so it shares the play buckets), before the retryable? and permission checks. The shared check, short-circuit and log entry live in Ci::JobRateLimitable; each service keeps its own literal feature flag check. The per-job check runs first and short-circuits the per-project check, so a caller already blocked on one job does not consume the project-wide budget. Throttled calls are logged with Gitlab::AppJsonLogger.

Exempted with rate_limit: false, because they are system-started:

  • Ci::Build auto-retry after a failure.
  • Deployments::AutoRollbackService (EE).
  • The retry fallback inside Ci::PlayBuildService, so a play is not also counted as a retry.
  • Environment stop actions (Environment#run_stop_action!, reached from auto-stop, merge request close and branch deletion as well as a manual stop), so a throttled on_stop job cannot leave an environment stopped without running it. Ci::Build#play passes rate_limit: through for this.

Two other callers of Ci::Build#play keep counting but no longer swallow a rejected play: Ci::PlayManualStageService logs it like an access error, and the ChatOps /deploy command answers with the reason instead of failing on a missing job.

Ci::RetryPipelineService uses clone!, not execute, and is not affected.

How it surfaces:

  • REST: POST /projects/:id/jobs/:job_id/retry and POST /projects/:id/jobs/:job_id/play return 429 Too Many Requests with a Retry-After header taken from the rule period.
  • GraphQL: jobRetry and jobPlay return the message in errors.
  • Web UI: Projects::JobsController#retry and #play return 429 JSON for the pipeline graph action button, or a flash alert plus a redirect back to the job for the HTML flow.

Note that the controller and the REST endpoint reject a play of a job that is no longer playable? before the service runs, so a repeated play of the same job is stopped there with 400/422 rather than counted. The per-job play limit therefore mostly applies through GraphQL and the play fallback; the per-project limit covers the rest.

Both checks are behind gitlab_com_derisk feature flags, default off, milestone 19.5: rate_limit_job_retry (rollout #630118) and rate_limit_job_play (rollout #630119). Two flags so that play, which has firmer numbers, can roll out first.

Two admin settings are added under Admin > Settings > CI/CD > Continuous Integration and Deployment: "Maximum job retries per project" and "Maximum manual job runs per project". The Settings API and OpenAPI document are updated. Docs: doc/administration/cicd/limits.md (new section), doc/api/jobs.md (429), doc/api/settings.md, doc/user/gitlab_com/_index.md.

The same shape is used in !255403 (merged) (pipeline cancel) and !254367 (merged) (pipeline retry). Adjacent additions in supported_rate_limits.rb and labkit_adapter.rb mean a small merge conflict with those MRs is expected.

References

Screenshots or screen recordings

The only UI change is two new number fields on the admin CI/CD settings page.

Before After

How to set up and validate locally

  1. In a Rails console, enable a flag: Feature.enable(:rate_limit_job_play) (or rate_limit_job_retry).
  2. Lower the limit: ApplicationSetting.current.update!(job_play_limit_per_user_project: 1), or in Admin > Settings > CI/CD.
  3. Run two manual jobs in the same project through the UI or POST /api/v4/projects/:id/jobs/:job_id/play.
  4. The second call is rejected: 429 from the API, or a flash alert in the UI.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Daniel Prause

Merge request reports

Loading
Loading