Roll out MCP Server plan throttles in observe mode
Summary
Turns on log-only counting for the per-plan MCP Server rate limits on POST /api/v4/mcp, so normal traffic is visible before any request gets rejected. First of three issues; enforcement follows in #631324 and #631325. Tracking issue #629883 (closed), epic gitlab-org&22800, milestone 19.5.
- DRI:
@terrichu - Team Slack channel:
#g_agent_execution
What the flags do
| Flag | Effect when enabled |
|---|---|
rate_limiter_mcp_limits_free_info |
burst_mcp_traffic_per_user_free_plan_log starts matching: Free-plan MCP traffic gets counted, never rejected |
rate_limiter_mcp_limits_premium_info |
burst_mcp_traffic_per_user_premium_plan_log starts matching: Premium-plan MCP traffic gets counted, never rejected |
rate_limiter_mcp_limits_ultimate_info |
burst_mcp_traffic_per_user_ultimate_plan_log starts matching: Ultimate-plan MCP traffic gets counted, never rejected |
The three _enforce flags stay off. No MCP request gets rejected as part of this issue.
Prerequisites
- !257035 (merged) and !257085 (merged) merged and deployed to gprd. Merged is not deployed: check the deploy before flipping any flag in this issue.
What could go wrong
- Plan resolution now runs for MCP traffic. It reads
Rails.cache(Gitlab::Redis::Cache) underplan_tierterrichu:<id>, cached 10 minutes, Free on a miss. A cache miss costs one database query. This is the first step in the rollout that adds this read to the MCP request path. - Each
_logrule that matches adds one rate-limiting Redis key per requester per minute, underlabkit:rl:{rack_request_mcp:<rule name>:requester_type:<type>:requester_id:<id>}. - Labkit fails open: if rule evaluation errors, the request still goes through and the error gets counted instead of the request.
- The
:limitrules (burst_mcp_traffic_per_user_<plan>_plan) must never match in this issue. If their series shows up in the counters, something enabled_enforceahead of schedule.
Rollout
Per incremental rollout process, run these in #production and cross-post to the team channel. Ramp each _info flag with --actors, the same steps the rate limiting team used on the rate_limiter_plan_limits_*_info flags in gitlab-com/gl-infra/production-engineering#29687. The actor is Feature.current_request, so a percentage samples requests rather than users: the ramp is there to bring the plan-resolution cache reads and the new counter writes up gradually, not to give a partial picture. The _log counters are only complete at 100%, so read them there. Wait at least 15 minutes between steps and check the Redis and database panels.
Non-production:
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info true --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info true --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info true --dev --pre --staging --staging-refProduction, one flag at a time, Free first:
-
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info 5 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info 25 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info 50 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info 100 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info 50 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info 100 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info 50 --actors -
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info 100 --actors
Verify
Grafana: the rate-limiting detail dashboard, Application::Labkit Rate Limiter row. The tier-aware panels added by gitlab-com/runbooks!11499 (merged) read rate_limiter="rack_request" with rule=~".*_traffic_per_.*", so the MCP rules show up as their own series with no dashboard change.
The _log rules are live:
sum by (rule, result) (sli_aggregations:gitlab_labkit_rate_limiter_rule_evaluations_total:rate_5m{env="gprd", rate_limiter="rack_request", rule=~"burst_mcp_traffic_per_user_.*_log"})allow must appear once a flag is on. log is the would-be rejection rate, the number this issue exists to learn.
Nothing enforces:
sum(sli_aggregations:gitlab_labkit_rate_limiter_rule_evaluations_total:rate_5m{env="gprd", rule=~"burst_mcp_traffic_per_user_.*_plan"})Must have no series. If one appears, an _enforce flag went on ahead of schedule. gitlab_labkit_rate_limiter_enforced_total does not exist in gprd, so this check stays on rule_evaluations_total.
Fail-open stays flat:
sum(gitlab_component_errors:rate_5m{env="gprd", type="rate-limiting", component="rate_limiter_checks"})Kibana gives per-request attribution the counters cannot. On pubsub-rails-inf-gprd*, json.rate_limit_state carries one <limiter>:<rule>:<result> string per counted rule:
json.path:"/api/v4/mcp" and json.rate_limit_state:*burst_mcp*Roll back if: the fail-open rate rises, redis-cluster-ratelimiting saturation moves, or any burst_mcp_traffic_per_user_.*_plan series appears.
Two caveats when reading the log share. It is a sawtooth rather than one number, because every counter starts at the flip and expires together, so report a range and a mean. And every step below 100 percent understates the real share: at 5 percent of actors a requester needs roughly 20 times the traffic to reach the counted limit.
A flag flip does not reach every Puma worker at once, so give it a minute before reading counters. Observed on a GDK: straight after disabling an enforce flag, a 65-request run split 30 on the enforcing rule and 35 on the _log twin; re-run once the flag cache had settled it was 65 on the twin and nothing on the enforcing rule. A mixed first minute after each flip is expected, not a bug.
Runbook: Tier-aware throttles (labkit).
Rollback
Production:
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info false
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info false
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info falseNon-production:
/chatops gitlab run feature set rate_limiter_mcp_limits_free_info false --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_premium_info false --dev --pre --staging --staging-ref
/chatops gitlab run feature set rate_limiter_mcp_limits_ultimate_info false --dev --pre --staging --staging-refCleanup
Remove per cleaning up:
/chatops gitlab run feature delete rate_limiter_mcp_limits_free_info --dev --pre --staging --staging-ref --production
/chatops gitlab run feature delete rate_limiter_mcp_limits_premium_info --dev --pre --staging --staging-ref --production
/chatops gitlab run feature delete rate_limiter_mcp_limits_ultimate_info --dev --pre --staging --staging-ref --productionHold off on running these until #631324 and #631325 are stable: the :limit rules need _info to stay true alongside _enforce, so deleting _info early would break enforcement. The actual YAML removal happens together with the _enforce flags in one cleanup MR, tracked on #631325.