Enforce Premium and Ultimate plan MCP Server throttles
Summary
Switches Premium- and Ultimate-plan MCP Server traffic from log-only counting to enforcement: requests over 600 a minute get rejected with a 429. Last of three issues; follows #631323 and #631324. Tracking issue #629883 (closed), epic gitlab-org&22800, milestone 19.5.
- DRI:
@terrichu - Team Slack channel:
#g_agent_execution
What the flags do
| Flag | Effect when enabled |
|---|---|
rate_limiter_mcp_limits_premium_enforce |
burst_mcp_traffic_per_user_premium_plan (the :limit rule) starts matching instead of its _log twin: Premium-plan MCP traffic over 600 a minute gets rejected with a 429 |
rate_limiter_mcp_limits_ultimate_enforce |
burst_mcp_traffic_per_user_ultimate_plan (the :limit rule) starts matching instead of its _log twin: Ultimate-plan MCP traffic over 600 a minute gets rejected with a 429 |
Free-plan enforcement is already live from #631324 by the time this issue runs.
Prerequisites
- #631323 completed and the
burst_mcp_traffic_per_user_premium_plan_logandburst_mcp_traffic_per_user_ultimate_plan_logcounters reviewed, so the 600-a-minute limit is known not to reject normal traffic on either plan -
rate_limiter_mcp_limits_premium_infoandrate_limiter_mcp_limits_ultimate_infostill enabled (from #631323); the:limitrules need them true alongside_enforce
What could go wrong
- Matching switches from the
_logrule's counter key to the:limitrule's counter key for each plan enabled; still one rate-limiting Redis key per requester per minute, same as during observe. - Labkit fails open: if rule evaluation errors, the request still goes through.
- Premium and Ultimate users over 600 MCP requests a minute start getting 429s. 600 is looser than the Free plan's 60, but still worth confirming against the observe-mode counters plan by plan.
Rollout
Per incremental rollout process, run in #production and cross-post to the team channel. Enable one plan, watch it, then enable the other. No percentage step here: the actor is Feature.current_request, so a partial flip would reject a user on some requests and let the same user through on others, at the same request rate. Enforcement is only coherent at 0% or 100%. Each flag goes straight from off to fully on.
Non-production:
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_enforce true --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_enforce true --dev --pre --staging --staging-refProduction, one plan at a time:
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_enforce trueWatch the burst_mcp_traffic_per_user_premium_plan series and 429 rate before moving on.
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_enforce trueVerify
Grafana: the rate-limiting detail dashboard, Application::Labkit Rate Limiter row. The tier-aware panels added by gitlab-com/runbooks!11499 (merged) read rate_limiter="rack_request" with rule=~".*_traffic_per_.*", so the MCP rules show up as their own series with no dashboard change.
Enforcement is live, and result="block" is the 429 count. There is no dedicated per-throttle 429 counter, and gitlab_labkit_rate_limiter_enforced_total does not exist in gprd.
sum by (rule, result) (sli_aggregations:gitlab_labkit_rate_limiter_rule_evaluations_total:rate_5m{env="gprd", rate_limiter="rack_request", rule=~"burst_mcp_traffic_per_user_(premium|ultimate)_plan"})Expect block to appear for whichever plan was just flipped, and its _log twin to fall away. The block rate should sit at or below the log rate measured in #631323 — materially higher means the flip changed which requests match, not just what happens to them.
Fail-open stays flat:
sum(gitlab_component_errors:rate_5m{env="gprd", type="rate-limiting", component="rate_limiter_checks"})Kibana: on pubsub-rails-inf-gprd*, json.rate_limit_state carries one <limiter>:<rule>:<result> string per counted rule.
json.path:"/api/v4/mcp" and json.rate_limit_state:*burst_mcp*A rejected request never reaches Grape, so a 429 produces no rate_limit_state entry. Use Kibana for what still gets through and Prometheus for what is blocked.
Roll back if: the fail-open rate rises, redis-cluster-ratelimiting saturation moves, or the block rate exceeds the log-mode share from #631323.
A flag flip does not reach every Puma worker at once, so give it a minute before reading counters. Observed on a GDK: straight after disabling an enforce flag, a 65-request run split 30 on the enforcing rule and 35 on the _log twin; re-run once the flag cache had settled it was 65 on the twin and nothing on the enforcing rule. A mixed first minute after each flip is expected, not a bug.
Runbook: Tier-aware throttles (labkit).
Rollback
Production:
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_enforce false
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_enforce falseNon-production:
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_enforce false --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_enforce false --dev --pre --staging --staging-refCleanup
Remove per cleaning up:
/chatops gitlab run feature delete rate_limiter_mcp_limits_premium_enforce --dev --pre --staging --staging-ref --production
/chatops gitlab run feature delete rate_limiter_mcp_limits_ultimate_enforce --dev --pre --staging --staging-ref --productionThis is the last of the three rollout issues. Delete all six MCP flags together in one cleanup MR: the three _info flags, the three _enforce flags, their YAML definitions, and the FlagPolicy predicates that reference them.