Add per-plan MCP Server rate limits for GitLab.com
What does this MR do and why?
Adds per-plan rate limits for the MCP Server on GitLab.com: 60 requests a minute for a Free
requester, 600 for Premium and Ultimate, counted per requester. Each plan gets an enforcing
rule and a log-only twin, switched by its own two gitlab_com_derisk flags, so a tier can be
watched before it rejects anything.
Self-managed and Dedicated are unaffected. No plan rule can match off GitLab.com, so the
throttle_authenticated_mcp_enabled setting from !257035 (merged) stays their MCP limit.
References
- Issue: #629883 (closed)
- Epic: &22800
- Rollout: #631323 (observe), #631324 (enforce Free), #631325 (enforce Premium and Ultimate)
- Depends on: !257035 (merged)
- Related: !258066 (merged), which removes
ratelimiting_include_plan_info. These rules matchrequester_plan, which that flag gates until it merges.
🤖 Decisions, for reviewers who want them
The rules match the tier-aware requester_plan fact. Earlier revisions resolved the plan
into a fact of their own, so the limits would keep working if ratelimiting_include_plan_info
were off — @bastirehm had asked for plan detection we own. @ashs2 pointed out that flag is
being removed, after which identity_facts emits requester_plan on every GitLab.com request,
so a second fact was resolving the same value under another name. That converges with the
approach in !257072 (closed), which is closed: the difference is that the flag is going away rather than
being depended on.
That removal is !258066 (merged), approved and set to merge. Until it does, these rules need
ratelimiting_include_plan_info on, which it is on GitLab.com. Once it merges, the flag and the
plan_facts_enabled? check go away and requester_plan is gated on GitLab.com alone.
Two flags per plan. _info lets that plan's _log rule count; _enforce swaps it for the
rejecting rule, and only while _info is on. Same shape as the existing per-plan rules, as
@ashs2 asked for, and it lets one plan roll back without the other two.
The rules do not match throttle_authenticated_mcp_enabled. It defaults to false and is off
on GitLab.com, so matching it would leave every flag inert until an admin turned it on. Without
it the six flags are the whole rollout: no change request, no dry-run variable. The consequence
worth knowing: disabling that setting does not stop tier 429s on GitLab.com. The lever is
/chatops run feature set rate_limiter_mcp_limits_<plan>_enforce false.
The rules sit on the general limiter, not their own. They lead the rack_request list, so
they are evaluated before the general plan rules. They carry a path match, so they still only
fire on MCP requests, and plan rules never claim, so nothing below them stops counting. The flat
throttle_authenticated_mcp throttle from !257035 (merged) keeps its own rack_request_mcp limiter, so
an MCP request can still count against throttle_authenticated_api. Moving that throttle too
would end that, which is a separate decision and not taken here.
These limits are per-minute only. The hourly ceiling is the existing
sustained_traffic_per_user_<plan>_plan rule, which counts the same requests. So a Premium user
at 600/min reaches the 15,000/hour plan ceiling in roughly 25 minutes, Ultimate at 42.
@reprazent raised that on &22800; @amandarueda approved the numbers for GA with a
post-GA follow-up on long-running low-HITL flows, and @bastirehm agreed.
Screenshots or screen recordings
No UI changes.
How to set up and validate locally
- Put the GDK in SaaS mode (
GITLAB_SIMULATE_SAAS=1) and restartrails-web. Without itPlanRules.for_limiter('rack_request')returns no MCP rule. - Leave the throttle setting off and turn the six flags on, in
rails console. The plan fact the rules match still needsratelimiting_include_plan_infountil its removal merges:ApplicationSetting.current.update!(throttle_authenticated_mcp_enabled: false) Feature.enable(:ratelimiting_include_plan_info) # not needed once !258066 merges %i[free premium ultimate].each do |p| Feature.enable(:"rate_limiter_mcp_limits_#{p}_info") Feature.enable(:"rate_limiter_mcp_limits_#{p}_enforce") end - Check which tier each user resolves to:
Gitlab::RateLimit::TierResolver.tier_for('user', User.find_by_username('root').id) - Mint an OAuth token with the
mcpscope for a Free user and a paid one. That is what MCP clients send; a personal access token cannot carry that scope outside the admin API. POST /api/v4/mcpwithAuthorization: Bearer <token>in a loop, past that tier's limit. Expect429at 61 for Free and 601 for paid, withRateLimit-Namenaming the rule.- Turn one plan's
_enforceflag off and repeat. That plan should count in its_logtwin and never reject. Wait about a minute after a flag flip before reading counters.
🤖 What I observed
GDK at 856b8b32a119, SaaS mode, throttle_authenticated_mcp_enabled off,
ratelimiting_include_plan_info on, OAuth tokens carrying the mcp scope.
| Case | Result |
|---|---|
| Free, 61 requests | 60 × 200, then 429. RateLimit-Name: burst_mcp_traffic_per_user_free_plan, RateLimit-Limit: 60, RateLimit-Observed: 62, Retry-After: 32 |
| Redis after that run | one MCP key, labkit:rl:{rack_request:burst_mcp_traffic_per_user_free_plan:requester_type:user:requester_id:3}. On the general limiter, and no authenticated_mcp key, so with the setting off the plan rule counts alone |
| Ultimate, same 61 requests | 61 × 200, so the Free ceiling does not apply |
| Ultimate seeded to 600 | 429, RateLimit-Name: burst_mcp_traffic_per_user_ultimate_plan, RateLimit-Limit: 600 |
Free with _enforce off |
65 × 200, all 65 on burst_mcp_traffic_per_user_free_plan_log, nothing on the enforcing rule |
A flag flip is not visible to every Puma worker at once, so counters should be read about a minute after a flip rather than immediately.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.