Add per-plan MCP Server rate limits for GitLab.com

What does this MR do and why?

Adds per-plan rate limits for the MCP Server on GitLab.com: 60 requests a minute for a Free requester, 600 for Premium and Ultimate, counted per requester. Each plan gets an enforcing rule and a log-only twin, switched by its own two gitlab_com_derisk flags, so a tier can be watched before it rejects anything.

Self-managed and Dedicated are unaffected. No plan rule can match off GitLab.com, so the throttle_authenticated_mcp_enabled setting from !257035 (merged) stays their MCP limit.

References

🤖 Decisions, for reviewers who want them

The rules match the tier-aware requester_plan fact. Earlier revisions resolved the plan into a fact of their own, so the limits would keep working if ratelimiting_include_plan_info were off — @bastirehm had asked for plan detection we own. @ashs2 pointed out that flag is being removed, after which identity_facts emits requester_plan on every GitLab.com request, so a second fact was resolving the same value under another name. That converges with the approach in !257072 (closed), which is closed: the difference is that the flag is going away rather than being depended on.

That removal is !258066 (merged), approved and set to merge. Until it does, these rules need ratelimiting_include_plan_info on, which it is on GitLab.com. Once it merges, the flag and the plan_facts_enabled? check go away and requester_plan is gated on GitLab.com alone.

Two flags per plan. _info lets that plan's _log rule count; _enforce swaps it for the rejecting rule, and only while _info is on. Same shape as the existing per-plan rules, as @ashs2 asked for, and it lets one plan roll back without the other two.

The rules do not match throttle_authenticated_mcp_enabled. It defaults to false and is off on GitLab.com, so matching it would leave every flag inert until an admin turned it on. Without it the six flags are the whole rollout: no change request, no dry-run variable. The consequence worth knowing: disabling that setting does not stop tier 429s on GitLab.com. The lever is /chatops run feature set rate_limiter_mcp_limits_<plan>_enforce false.

The rules sit on the general limiter, not their own. They lead the rack_request list, so they are evaluated before the general plan rules. They carry a path match, so they still only fire on MCP requests, and plan rules never claim, so nothing below them stops counting. The flat throttle_authenticated_mcp throttle from !257035 (merged) keeps its own rack_request_mcp limiter, so an MCP request can still count against throttle_authenticated_api. Moving that throttle too would end that, which is a separate decision and not taken here.

These limits are per-minute only. The hourly ceiling is the existing sustained_traffic_per_user_<plan>_plan rule, which counts the same requests. So a Premium user at 600/min reaches the 15,000/hour plan ceiling in roughly 25 minutes, Ultimate at 42. @reprazent raised that on &22800; @amandarueda approved the numbers for GA with a post-GA follow-up on long-running low-HITL flows, and @bastirehm agreed.

Screenshots or screen recordings

No UI changes.

How to set up and validate locally

  1. Put the GDK in SaaS mode (GITLAB_SIMULATE_SAAS=1) and restart rails-web. Without it PlanRules.for_limiter('rack_request') returns no MCP rule.
  2. Leave the throttle setting off and turn the six flags on, in rails console. The plan fact the rules match still needs ratelimiting_include_plan_info until its removal merges:
    ApplicationSetting.current.update!(throttle_authenticated_mcp_enabled: false)
    Feature.enable(:ratelimiting_include_plan_info) # not needed once !258066 merges
    %i[free premium ultimate].each do |p|
      Feature.enable(:"rate_limiter_mcp_limits_#{p}_info")
      Feature.enable(:"rate_limiter_mcp_limits_#{p}_enforce")
    end
  3. Check which tier each user resolves to:
    Gitlab::RateLimit::TierResolver.tier_for('user', User.find_by_username('root').id)
  4. Mint an OAuth token with the mcp scope for a Free user and a paid one. That is what MCP clients send; a personal access token cannot carry that scope outside the admin API.
  5. POST /api/v4/mcp with Authorization: Bearer <token> in a loop, past that tier's limit. Expect 429 at 61 for Free and 601 for paid, with RateLimit-Name naming the rule.
  6. Turn one plan's _enforce flag off and repeat. That plan should count in its _log twin and never reject. Wait about a minute after a flag flip before reading counters.
🤖 What I observed

GDK at 856b8b32a119, SaaS mode, throttle_authenticated_mcp_enabled off, ratelimiting_include_plan_info on, OAuth tokens carrying the mcp scope.

Case Result
Free, 61 requests 60 × 200, then 429. RateLimit-Name: burst_mcp_traffic_per_user_free_plan, RateLimit-Limit: 60, RateLimit-Observed: 62, Retry-After: 32
Redis after that run one MCP key, labkit:rl:{rack_request:burst_mcp_traffic_per_user_free_plan:requester_type:user:requester_id:3}. On the general limiter, and no authenticated_mcp key, so with the setting off the plan rule counts alone
Ultimate, same 61 requests 61 × 200, so the Free ceiling does not apply
Ultimate seeded to 600 429, RateLimit-Name: burst_mcp_traffic_per_user_ultimate_plan, RateLimit-Limit: 600
Free with _enforce off 65 × 200, all 65 on burst_mcp_traffic_per_user_free_plan_log, nothing on the enforcing rule

A flag flip is not visible to every Puma worker at once, so counters should be read about a minute after a flip rather than immediately.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Terri Chu

Merge request reports

Loading
Loading