Roll out MCP Server plan throttles in observe mode

Summary

Turns on log-only counting for the per-plan MCP Server rate limits on POST /api/v4/mcp, so normal traffic is visible before any request gets rejected. First of three issues; enforcement follows in #631324 and #631325. Tracking issue #629883 (closed), epic gitlab-org&22800, milestone 19.5.

  • DRI: @terrichu
  • Team Slack channel: #g_agent_execution

What the flags do

Flag Effect when enabled
rate_limiter_mcp_limits_free_info burst_mcp_traffic_per_user_free_plan_log starts matching: Free-plan MCP traffic gets counted, never rejected
rate_limiter_mcp_limits_premium_info burst_mcp_traffic_per_user_premium_plan_log starts matching: Premium-plan MCP traffic gets counted, never rejected
rate_limiter_mcp_limits_ultimate_info burst_mcp_traffic_per_user_ultimate_plan_log starts matching: Ultimate-plan MCP traffic gets counted, never rejected

The three _enforce flags stay off. No MCP request gets rejected as part of this issue.

Prerequisites

What could go wrong

  • Plan resolution now runs for MCP traffic. It reads Rails.cache (Gitlab::Redis::Cache) under plan_tierterrichu:<id>, cached 10 minutes, Free on a miss. A cache miss costs one database query. This is the first step in the rollout that adds this read to the MCP request path.
  • Each _log rule that matches adds one rate-limiting Redis key per requester per minute, under labkit:rl:{rack_request_mcp:<rule name>:requester_type:<type>:requester_id:<id>}.
  • Labkit fails open: if rule evaluation errors, the request still goes through and the error gets counted instead of the request.
  • The :limit rules (burst_mcp_traffic_per_user_<plan>_plan) must never match in this issue. If their series shows up in the counters, something enabled _enforce ahead of schedule.

Rollout

Per incremental rollout process, run these in #production and cross-post to the team channel. Ramp each _info flag with --actors, the same steps the rate limiting team used on the rate_limiter_plan_limits_*_info flags in gitlab-com/gl-infra/production-engineering#29687. The actor is Feature.current_request, so a percentage samples requests rather than users: the ramp is there to bring the plan-resolution cache reads and the new counter writes up gradually, not to give a partial picture. The _log counters are only complete at 100%, so read them there. Wait at least 15 minutes between steps and check the Redis and database panels.

Non-production:

/chatops gitlab run feature set rate_limiter_mcp_limits_free_info true --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info true --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info true --dev --pre --staging --staging-ref

Production, one flag at a time, Free first:

  • /chatops gitlab run feature set rate_limiter_mcp_limits_free_info 5 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_free_info 25 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_free_info 50 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_free_info 100 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_premium_info 50 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_premium_info 100 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info 50 --actors
  • /chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info 100 --actors

Verify

Grafana: the rate-limiting detail dashboard, Application::Labkit Rate Limiter row. The tier-aware panels added by gitlab-com/runbooks!11499 (merged) read rate_limiter="rack_request" with rule=~".*_traffic_per_.*", so the MCP rules show up as their own series with no dashboard change.

The _log rules are live:

sum by (rule, result) (sli_aggregations:gitlab_labkit_rate_limiter_rule_evaluations_total:rate_5m{env="gprd", rate_limiter="rack_request", rule=~"burst_mcp_traffic_per_user_.*_log"})

allow must appear once a flag is on. log is the would-be rejection rate, the number this issue exists to learn.

Nothing enforces:

sum(sli_aggregations:gitlab_labkit_rate_limiter_rule_evaluations_total:rate_5m{env="gprd", rule=~"burst_mcp_traffic_per_user_.*_plan"})

Must have no series. If one appears, an _enforce flag went on ahead of schedule. gitlab_labkit_rate_limiter_enforced_total does not exist in gprd, so this check stays on rule_evaluations_total.

Fail-open stays flat:

sum(gitlab_component_errors:rate_5m{env="gprd", type="rate-limiting", component="rate_limiter_checks"})

Kibana gives per-request attribution the counters cannot. On pubsub-rails-inf-gprd*, json.rate_limit_state carries one <limiter>:<rule>:<result> string per counted rule:

json.path:"/api/v4/mcp" and json.rate_limit_state:*burst_mcp*

Roll back if: the fail-open rate rises, redis-cluster-ratelimiting saturation moves, or any burst_mcp_traffic_per_user_.*_plan series appears.

Two caveats when reading the log share. It is a sawtooth rather than one number, because every counter starts at the flip and expires together, so report a range and a mean. And every step below 100 percent understates the real share: at 5 percent of actors a requester needs roughly 20 times the traffic to reach the counted limit.

A flag flip does not reach every Puma worker at once, so give it a minute before reading counters. Observed on a GDK: straight after disabling an enforce flag, a 65-request run split 30 on the enforcing rule and 35 on the _log twin; re-run once the flag cache had settled it was 65 on the twin and nothing on the enforcing rule. A mixed first minute after each flip is expected, not a bug.

Runbook: Tier-aware throttles (labkit).

Rollback

Production:

/chatops gitlab run feature set rate_limiter_mcp_limits_free_info false
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info false
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info false

Non-production:

/chatops gitlab run feature set rate_limiter_mcp_limits_free_info false --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info false --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info false --dev --pre --staging --staging-ref

Cleanup

Remove per cleaning up:

/chatops gitlab run feature delete rate_limiter_mcp_limits_free_info --dev --pre --staging --staging-ref --production
/chatops gitlab run feature delete rate_limiter_mcp_limits_premium_info --dev --pre --staging --staging-ref --production
/chatops gitlab run feature delete rate_limiter_mcp_limits_ultimate_info --dev --pre --staging --staging-ref --production

Hold off on running these until #631324 and #631325 are stable: the :limit rules need _info to stay true alongside _enforce, so deleting _info early would break enforcement. The actual YAML removal happens together with the _enforce flags in one cleanup MR, tracked on #631325.

Edited by Terri Chu