Enforce the MCP Server rate limit at the request edge

What does this MR do and why?

Adds the Labkit rate-limit rule and the Rack::Attack throttle definition that enforce the MCP settings added in !256468 (merged). POST /api/v4/mcp is now counted per requester against throttle_authenticated_mcp_*, in its own limiter (rack_request_mcp) so it keeps counting against the general authenticated API throttle rather than claiming those requests away from it.

Gated by the throttle_authenticated_mcp_enabled setting alone, which ships false.

🤖 Why there is no feature flag

An earlier revision added rate_limit_mcp_server, a gitlab_com_derisk flag, and it has been removed. What it offered over the setting was a percentage rollout on GitLab.com. Everything else it would have done is already covered:

  • Observe mode comes from GITLAB_THROTTLE_DRY_RUN. This is a registry throttle, so Limiters#build_rules builds it as :log when the throttle is named in that variable — it counts and logs, and rejects nothing.
  • An instant off switch is the setting itself, changeable through the settings API with no deploy.
  • Per-plan rollout and rollback belong to the per-plan limits in !257085 (merged), which carry their own info and enforce flags per tier.

Against that, a gitlab_com_derisk flag does not exist on self-managed or Dedicated, so while it existed an administrator who enabled the setting there would get nothing — making flag removal a GA blocker we would have created for ourselves. Discussed with @ashs2 and @reprazent on &22800.

The rollout is therefore: name the throttle in GITLAB_THROTTLE_DRY_RUN, enable the setting, measure, then remove the name so it enforces. Tracked in #630542 (closed).

Requires !256468 (merged), now merged.

🤖 Design notes
  • Own limiter, not the general one. A registry entry with claims: true terminates evaluation for the requests it matches, so in a shared limiter the MCP rule and the throttle_authenticated_api rule could not both count the same request. rack_request_protected_paths exists for the same reason. (Per-plan rules are unaffected either way: Limiters#build puts PlanRules.for_limiter ahead of every registry rule.)
  • Counts every MCP request, not only tools/call. The middleware runs above Rack::Attack, before routing, and cannot read the JSON-RPC body. Splitting by method would need enforcement inside the Grape endpoint.
  • No unauthenticated counterpart. authenticate! is the first line of the before block on the whole MCP namespace, so an unauthenticated request never reaches a countable state; the rule's requester_id: /./ gate reflects that.
  • MCP_PATH_REGEX is ^/api/v\d+/mcp(?:/|$). Anchored at both ends, so /api/v4/orbit/mcp, the .well-known OAuth discovery paths and any future /api/v4/mcp_* route stay out of this throttle.

References

Screenshots or screen recordings

No UI changes.

How to set up and validate locally

Validated against a GDK running this branch. No feature flags to touch: rate_limiter_use_labkit_rack_cohort_1 and its _enforce twin already default to on.

  1. Enable the throttle with a low limit so the boundary is reachable, and mint a token. In rails console:

    ApplicationSetting.current.update!(
      throttle_authenticated_mcp_enabled: true,
      throttle_authenticated_mcp_requests_per_period: 3,
      throttle_authenticated_mcp_period_in_seconds: 60
    )
    user = User.find_by_username('root')
    token = user.personal_access_tokens.create!(
      name: 'mcp-limit-test', scopes: [:api, :mcp], expires_at: 30.days.from_now
    )
    token.token
  2. Send five requests to the MCP endpoint with that token:

    for i in 1 2 3 4 5; do
      curl -ks -o /dev/null -D - -X POST "https://gdk.test:3443/api/v4/mcp" \
        -H "PRIVATE-TOKEN: $PAT" -H 'Content-Type: application/json' \
        -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}' \
        | grep -iE '^(HTTP/|ratelimit-|retry-after)'
    done

    The first three return 200, the fourth and fifth 429:

    req 1 -> 200  ratelimit-limit=3 ratelimit-name=throttle_authenticated_mcp ratelimit-observed=1 ratelimit-remaining=2
    req 2 -> 200  ratelimit-limit=3 ratelimit-name=throttle_authenticated_mcp ratelimit-observed=2 ratelimit-remaining=1
    req 3 -> 200  ratelimit-limit=3 ratelimit-name=throttle_authenticated_mcp ratelimit-observed=3 ratelimit-remaining=0
    req 4 -> 429  ... ratelimit-resettime=Wed, 23 Sep 2026 12:03:00 GMT retry-after=53
    req 5 -> 429  ... ratelimit-resettime=Wed, 23 Sep 2026 12:03:00 GMT retry-after=53

    RateLimit-Limit, -Name, -Observed, -Remaining and -Reset are present on the 200s as well as the 429s, so a client can see the limit approaching. -ResetTime and Retry-After are added on rejection only.

  3. Confirm the count landed in the MCP limiter's own keyspace, keyed per requester:

    Gitlab::Redis::RateLimiting.with { |r| r.keys('labkit:rl:*mcp*') }
    # => ["labkit:rl:{rack_request_mcp:authenticated_mcp:requester_type:user:requester_id:1}"]
  4. Confirm one user exhausting the limit does not affect another. Mint a token for a second user and send four requests as them, then one more as the first user:

    user2 req 1: 200 observed=1 remaining=2
    user2 req 2: 200 observed=2 remaining=1
    user2 req 3: 200 observed=3 remaining=0
    user2 req 4: 429 observed=4
    user1 again: 200 observed=2 remaining=1

    Two independent counters, one per requester:

    labkit:rl:{rack_request_mcp:authenticated_mcp:requester_type:user:requester_id:1} = 2
    labkit:rl:{rack_request_mcp:authenticated_mcp:requester_type:user:requester_id:2} = 4
  5. Set throttle_authenticated_mcp_enabled: false and confirm the same loop no longer returns 429 and writes no counter.

  6. Clean up: disable the setting, restore the defaults (600 / 60), and revoke the test tokens.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Edited by Terri Chu

Merge request reports

Loading
Loading