Log the matched rate-limit rule and its result

What

Puts one field on the per-request log for requests that go through the Labkit rack middleware:

  • rate_limit_state, an array of "<limiter>:<rule>:<result>" strings

result is one of allow, log, block, banned, skip, the same values as the result label on rule_evaluations_total, so a log line and a metric series pivot on the same value. A boolean would have folded banned into block and dropped skip entirely.

One entry per rule that counted, so a request that trips a plan rule and a registry rule on the same limiter shows both. Only counting rules appear: the claim twins, the dry-run bypasses and the skip rules terminate the check rather than counting, so they never become an evaluation and drop out without being filtered by name. Where nothing counted at all, the rule that terminated the check is reported with result skip, so a bypassed request still says why it carries no other rule.

Three small pieces: a SafeRequestStore-backed Gitlab::Instrumentation::RateLimitState built the same way as RateLimitingGates, one call in the middleware after the results come back, and one merge in InstrumentationHelper.

Why

Today only ApplicationRateLimiter records anything, through its single call to RateLimitingGates.track. That is the only caller in the repo. The rack middleware records nothing, so when it throttles something there is no way to tell from the logs which rule did it.

That is most of the throttling we do. Over one window on 17 September the middleware blocked 3,393,527 requests against the application limiters' 770,113.

Enforcement starts in October, so this is the gap worth closing first.

Verified on real traffic

A local cluster in edit mode, two runs of 8 unauthenticated API calls against a limit of 5 per 60 seconds. Under the limit: ["rack_request:unauthenticated_api:allow"]. Over the limit with the throttle enforcing: HTTP 429 and no log line at all. Over the limit with GITLAB_THROTTLE_DRY_RUN set, which is the shape the observe window runs in: HTTP 200 and ["rack_request:unauthenticated_api:log"]. Detail in the comments.

What block actually means here, because it is not obvious

rate_limit_result: block means the rule went over its limit, not that the request was rejected. It can only ever appear on a request that succeeded.

Two separate reasons:

  • Registry throttles are built as :limit rules whatever their cohort (see the comment at the top of Limiters). The cohort decides whether a match turns into a 429, not whether the rule counts. So a shadow-cohort throttle going over limit logs block next to a 200.
  • When Labkit does reject a request, the middleware returns the 429 without calling the app. No controller runs, so nothing writes a log line at all.

The tier-aware plan rules behave differently, because their :limit variants carry plan_limits_<plan>_enforce: true in the match itself. With the flag off they never fire and you only see the _log twin.

Do not build a panel counting real blocks off this field. Getting a line out of a genuine Labkit 429 needs a dedicated log line from the middleware, which is separate work and is only needed once enforcement is on.

banned cannot occur here yet either. No rack registry rule and no plan rule sets ban_for, so four of the five values are reachable from this middleware today.

Resolved since the first version

One joined field rather than two arrays. Elasticsearch keeps no pairing between two multi-valued fields on a document, so rate_limit_rule:X AND rate_limit_result:Y matched a line where X and Y came from different entries, which is exactly the query a rollout dashboard runs. Asked for by the review on this merge request from three directions, and it closes the second question at the same time.

The limiter name is now recorded. It leads each entry, because every limiter builds the synthetic rules under the same name, so a rule name alone does not say which limiter reported it. This was also the original request on the issue: an array of limiter name plus rule name.

One thing I would still like a view on: it ships unsampled

Measured on gprd, 2026-09-21 11:30 to 12:00 UTC, the three rack limiters returned a matched result 75,368 times a second (rack_request 54,212, rack_request_protected_paths 15,687, rack_request_incident_management 5,468), against 55,327 Rails requests a second over 11:00 to 12:00 UTC.

sum by (rate_limiter) (rate(gitlab_labkit_rate_limiter_checks_total{env="gprd", matched="true"}[30m]))
avg_over_time(sum(rate(gitlab_transaction_duration_seconds_count{env="gprd"}[5m]))[1h:5m])

The field adds 22 bytes of key per line plus 34 to 54 bytes per entry, so 324 to 505 GB a day of raw log, depending on rule-name length. The estimate on the issue was 265 to 356 GB a day, which assumed one entry per request.

Nothing here samples. Emitting only log, block and banned, or 1 percent of allow and 100 percent of the rest, would cut it roughly sixfold. Happy to add it before this merges if that is the call.

Edited by Nidhey Indurkar

Merge request reports

Loading
Loading