Draft: Add per-plan MCP Server rate limits for GitLab.com
What does this MR do and why?
Adds a per-plan rate limit for the MCP endpoint on GitLab.com: 60 requests a
minute for a Free requester, 600 for Premium and Ultimate. They are labkit rules
on the rack_request_mcp limiter added in !257035 (merged), so they sit alongside the
instance-wide ceiling the throttle setting carries rather than replacing it.
Numbers from &22800. GitLab.com only, like every other plan rule.
Not GA scope, and not ready to merge — it is a concrete proposal to review
with the tier-aware rate limiting DRIs. See the open questions below.
@ashs2 @nindurkar @reprazent
🤖 How it fits the tier-aware design
- The Tier-Aware Request Throttles design
decouples the plan facts from the throttles precisely so that "any service
owner who wants a tier-specific rule of their own" can use them, with "which
plans a rule acts on" living in the rule's own
match. That is what this does. - The numbers are literals because the tier-aware rules are literals: published limits are a commitment, and Phase 2 of the unified rate limiting architecture moves all of them to labkit config at once.
MCP_RULESis a second list becausefor_limitergives each limiter its own ordered rule set.RULESgoes torack_request; putting the MCP rules there would count MCP requests a second time in the general keyspace.- Self-managed is excluded three ways, unchanged from the existing rules:
for_limiterreturns nothing off SaaS,requester_planis only emitted on SaaS, and CE ships only the stub.
Not GA scope
MCP Server GA (19.5) ships on !256468 (merged) and !257035 (merged) alone: one instance-wide
per-user limit from an application setting, working the same on GitLab.com,
self-managed and Dedicated. This MR cannot be GA scope, because every part of
the tier machinery it uses is behind gitlab_com_derisk flags owned by another
team, Premium and Ultimate enforcement is held until January 2027, and
self-managed and Dedicated read no feature flags at all.
It is here to agree the shape, not to merge on the GA timeline.
Open questions for the rate-limiting DRIs
-
These rules have no enforce gate, and that is the blocking one. The general plan rules pair each
:limitrule with a:logtwin, mutually exclusive onplan_limits_<plan>_enforce. These carry no equivalent, so two things follow. There is no observe period: we cannot measure how many Free users a 60/min limit would reject before it starts rejecting them. And because a plan rule has no registry entry, the middleware returns the 429 unconditionally, so the only thing that can switch these off isratelimiting_include_plan_info— which would take every tier-aware rule down with it. That is your kill switch, not ours, and we should not be holding it.Proposed fix: a
rate_limiter_mcp_limits_{info,enforce}pair and a:logtwin per rule, matching the shape the general rules use. Happy to write it that way if you agree. -
Do the MCP numbers sit sensibly under the plan budget? A Free user spending 60 requests a minute on MCP uses 3,600 of their 5,000 hourly allowance; 600 a minute on Premium exhausts 15,000 an hour in 25 minutes. So 600/min is a burst ceiling that the sustained budget overrules well inside the hour. That is fine if nobody publishes it as a sustained MCP allowance, which is a docs question for &22800.
-
Timing on paid tiers. Premium and Ultimate enforcement is held to January 2027 until customers can buy capacity. MCP is a new surface at GA rather than a reduction of an existing entitlement, but the call is yours.
-
setting_authenticated_mcpin the match.UNAUTHENTICATED_RULESdeliberately carries nosetting_*fact, with a spec pinning that a plan limit is a product commitment rather than an admin toggle. These rules take the opposite line so the instance switch governs both layers together. Which convention do you want? -
PlanRules.active?does not know about these. It reads only the eightFLAGS, so the MCP rules run only because the MCP throttle is cohort 1 andrate_limiter_use_labkit_rack_cohort_1is on. Fails closed, so not a defect, but it couples them to a Rack::Attack migration flag that is meant to be deleted.
The rules are gated by the rate_limit_mcp_server feature flag and the
throttle_authenticated_mcp_enabled setting, the same two switches as the flat limit
in !257035 (merged), so nothing here applies before that one does.
References
- Issue: #629883 (closed)
- Epic: &22800
- Depends on: !256468 (merged) (settings), !257035 (merged) (the MCP limiter and rule)
- Design: Tier-Aware Request Throttles, gitlab-com/gl-infra&2122
Screenshots or screen recordings
No UI changes.
How to set up and validate locally
- Check out this branch, which includes !256468 (merged) and !257035 (merged).
- In
rails console, put the instance on SaaS and enable the throttle:Gitlab::Saas.stub(:feature_available?).and_return(true) # or use a SaaS GDK ApplicationSetting.current.update!( throttle_authenticated_mcp_enabled: true, throttle_authenticated_mcp_requests_per_period: 10_000, throttle_authenticated_mcp_period_in_seconds: 60 ) Feature.enable(:ratelimiting_include_plan_info) Feature.enable(:rate_limiter_use_labkit_rack_cohort_1) Feature.enable(:rate_limiter_use_labkit_rack_cohort_1_enforce) Gitlab::RackAttack::LabkitRateLimit::Limiters.reset! - Seed the Free counter close to its ceiling so you do not need 60 requests:
user = User.find_by_username('root') key = "labkit:rl:{rack_request_mcp:mcp_traffic_per_user_free_plan:requester_type:user:requester_id:#{user.id}}" Gitlab::Redis::RateLimiting.with { |r| r.set(key, 60) } - Send one MCP request with that user's PAT and confirm a
429whoseRateLimit-Nameismcp_traffic_per_user_free_planand whoseRateLimit-Limitis60:curl -si -X POST "$GDK_URL/api/v4/mcp" -H "PRIVATE-TOKEN: $PAT" \ -H 'Content-Type: application/json' \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' | head -12 - Confirm the instance-wide rule still counts alongside it:
Gitlab::Redis::RateLimiting.with { |r| r.keys('labkit:rl:{rack_request_mcp:*') }
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.