Draft: Add per-plan MCP Server rate limits for GitLab.com

What does this MR do and why?

Adds a per-plan rate limit for the MCP endpoint on GitLab.com: 60 requests a minute for a Free requester, 600 for Premium and Ultimate. They are labkit rules on the rack_request_mcp limiter added in !257035 (merged), so they sit alongside the instance-wide ceiling the throttle setting carries rather than replacing it.

Numbers from &22800. GitLab.com only, like every other plan rule.

Not GA scope, and not ready to merge — it is a concrete proposal to review with the tier-aware rate limiting DRIs. See the open questions below. @ashs2 @nindurkar @reprazent

🤖 How it fits the tier-aware design
  • The Tier-Aware Request Throttles design decouples the plan facts from the throttles precisely so that "any service owner who wants a tier-specific rule of their own" can use them, with "which plans a rule acts on" living in the rule's own match. That is what this does.
  • The numbers are literals because the tier-aware rules are literals: published limits are a commitment, and Phase 2 of the unified rate limiting architecture moves all of them to labkit config at once.
  • MCP_RULES is a second list because for_limiter gives each limiter its own ordered rule set. RULES goes to rack_request; putting the MCP rules there would count MCP requests a second time in the general keyspace.
  • Self-managed is excluded three ways, unchanged from the existing rules: for_limiter returns nothing off SaaS, requester_plan is only emitted on SaaS, and CE ships only the stub.

Not GA scope

MCP Server GA (19.5) ships on !256468 (merged) and !257035 (merged) alone: one instance-wide per-user limit from an application setting, working the same on GitLab.com, self-managed and Dedicated. This MR cannot be GA scope, because every part of the tier machinery it uses is behind gitlab_com_derisk flags owned by another team, Premium and Ultimate enforcement is held until January 2027, and self-managed and Dedicated read no feature flags at all.

It is here to agree the shape, not to merge on the GA timeline.

Open questions for the rate-limiting DRIs

  1. These rules have no enforce gate, and that is the blocking one. The general plan rules pair each :limit rule with a :log twin, mutually exclusive on plan_limits_<plan>_enforce. These carry no equivalent, so two things follow. There is no observe period: we cannot measure how many Free users a 60/min limit would reject before it starts rejecting them. And because a plan rule has no registry entry, the middleware returns the 429 unconditionally, so the only thing that can switch these off is ratelimiting_include_plan_info — which would take every tier-aware rule down with it. That is your kill switch, not ours, and we should not be holding it.

    Proposed fix: a rate_limiter_mcp_limits_{info,enforce} pair and a :log twin per rule, matching the shape the general rules use. Happy to write it that way if you agree.

  2. Do the MCP numbers sit sensibly under the plan budget? A Free user spending 60 requests a minute on MCP uses 3,600 of their 5,000 hourly allowance; 600 a minute on Premium exhausts 15,000 an hour in 25 minutes. So 600/min is a burst ceiling that the sustained budget overrules well inside the hour. That is fine if nobody publishes it as a sustained MCP allowance, which is a docs question for &22800.

  3. Timing on paid tiers. Premium and Ultimate enforcement is held to January 2027 until customers can buy capacity. MCP is a new surface at GA rather than a reduction of an existing entitlement, but the call is yours.

  4. setting_authenticated_mcp in the match. UNAUTHENTICATED_RULES deliberately carries no setting_* fact, with a spec pinning that a plan limit is a product commitment rather than an admin toggle. These rules take the opposite line so the instance switch governs both layers together. Which convention do you want?

  5. PlanRules.active? does not know about these. It reads only the eight FLAGS, so the MCP rules run only because the MCP throttle is cohort 1 and rate_limiter_use_labkit_rack_cohort_1 is on. Fails closed, so not a defect, but it couples them to a Rack::Attack migration flag that is meant to be deleted.

The rules are gated by the rate_limit_mcp_server feature flag and the throttle_authenticated_mcp_enabled setting, the same two switches as the flat limit in !257035 (merged), so nothing here applies before that one does.

References

Screenshots or screen recordings

No UI changes.

How to set up and validate locally

  1. Check out this branch, which includes !256468 (merged) and !257035 (merged).
  2. In rails console, put the instance on SaaS and enable the throttle:
    Gitlab::Saas.stub(:feature_available?).and_return(true) # or use a SaaS GDK
    ApplicationSetting.current.update!(
      throttle_authenticated_mcp_enabled: true,
      throttle_authenticated_mcp_requests_per_period: 10_000,
      throttle_authenticated_mcp_period_in_seconds: 60
    )
    Feature.enable(:ratelimiting_include_plan_info)
    Feature.enable(:rate_limiter_use_labkit_rack_cohort_1)
    Feature.enable(:rate_limiter_use_labkit_rack_cohort_1_enforce)
    Gitlab::RackAttack::LabkitRateLimit::Limiters.reset!
  3. Seed the Free counter close to its ceiling so you do not need 60 requests:
    user = User.find_by_username('root')
    key = "labkit:rl:{rack_request_mcp:mcp_traffic_per_user_free_plan:requester_type:user:requester_id:#{user.id}}"
    Gitlab::Redis::RateLimiting.with { |r| r.set(key, 60) }
  4. Send one MCP request with that user's PAT and confirm a 429 whose RateLimit-Name is mcp_traffic_per_user_free_plan and whose RateLimit-Limit is 60:
    curl -si -X POST "$GDK_URL/api/v4/mcp" -H "PRIVATE-TOKEN: $PAT" \
      -H 'Content-Type: application/json' \
      -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -12
  5. Confirm the instance-wide rule still counts alongside it:
    Gitlab::Redis::RateLimiting.with { |r| r.keys('labkit:rl:{rack_request_mcp:*') }

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Terri Chu

Merge request reports

Loading
Loading