Reduce OpenBao requests when loading the secrets page

What does this MR do and why?

Loading /<project>/-/secrets (will) resolve the user's effective OpenBao capabilities twice: once in Projects::SecretsController#check_read_capability! to gate the page, and (eventually when implemented) once in the GraphQL userPermissions resolver. Those are separate requests, possibly on different Puma workers, so strong_memoize and SafeRequestStore can't share the result. Each resolution is three sequential OpenBao round trips: login, capabilities-self, and revoke.

This caches the OpenBao lookup in Rails.cache for 30 seconds, keyed per user and per secrets manager.

Benchmarks

ProjectEffectiveCapabilitiesService#execute on a GDK against a local OpenBao. n=20 per run, same project / secrets manager / user in all three, secrets_manager_paid_experience disabled throughout, cache key deleted before each cold iteration on the new code (measuring the cost of one lookup, independent of who calls it)

run min median p95 max
cache miss, before 23.45 28.04 36.70 43.93
cache miss, after 26.14 29.52 38.55 50.16
cache hit, after 0.12 0.12 0.29 0.29

The miss row is the control: unchanged within run-to-run noise, including the tail, so nothing expensive was added to the path that still calls OpenBao. A hit costs 0.12ms against ~28ms, so ~28ms is saved per lookup the cache absorbs.

OpenBao requests counted from the audit log across page loads:

OpenBao round trips, one page load before after, cold after, cached
capability lookup (3 round trips each) 3 3 0
(unchanged) secrets list — LIST detailed-metadata 2 2 2
total 5 5 2

Notes:

▎ Inline-authed requests appear twice in the audit log: the internal authentication, then the operation. So a page load shows 7 log entries for these 5 round trips.

▎ Today only the controller resolves capabilities, so the measured saving is on repeat views. Once the frontend requests userPermissions, a single view will resolve twice; the request spec covers the case inclusive of an assumed FE userPermisisons call, and confirms one lookup instead of two with the cache, which would take a cold cache view from 8 round trips (3 for capability lookup * 2 requests, plus 2 for the secrets list) to 5 (3 for the capability lookup once, plus 2 for the secrets list). The table only shows our results though, which currently does not include the userPermission request.

One capability lookup is three HTTP requests to OpenBao, in sequence:

  1. auth/user_jwt/cel/login — mint a real token (inline auth won't work, since capabilities-self resolves policies by looking the token up in the token store)
  2. sys/capabilities-self — the actual question
  3. auth/token/revoke-self — throw the token away

This is why the capability lookup row count is 3 before (would have been 6 if the FE userPermissions call landed before this).

Decisions worth a reviewer's attention

  • Entitlement is not cached. execute re-applies the read/write entitlement gates on every call. It changes independently of OpenBao (trial expiry, exhausted credits, lapsed subscription), so caching the combined booleans would freeze a billing decision.
  • Failures are not cached across requests. The fetch returns nil on the four OpenBao error classes and skip_nil: true keeps it out of Redis, so a transient blip can't deny a user for the whole TTL. Within a single request, Gitlab::Cache.fetch_once pins the nil in SafeRequestStore, so a second caller in the same request does not retry OpenBao unnecessarily.
  • Staleness is bounded at 30 seconds. A revoked or narrowed grant can keep working that long; Rails can't know when OpenBao policies change, so expiry is the mechanism. The GitLab policy layer is not cached, so a user removed from the project is still denied immediately.
  • Key uses secrets_manager.id. This also drops entries from a manager that was deprovisioned and replaced, and it doesn't assume each resource has only one secrets manager — project_secrets_managers enforces that with a unique index, but group_secrets_managers has no unique index on group_id. scope_name is in the key because the two manager tables have colliding IDs; current_user.id is required because the CEL program derives policies from the user.
  • No explicit invalidation when grants change. A grant attaches to a User, a Role, a Group, or a MemberRole. Only the single-user case maps to one cache key — a role or group grant affects every member holding it and can't be enumerated cheaply. Since those three would still depend on expiry, worst-case staleness is 30 seconds either way, so this leaves invalidation out rather than making revocation instant for one grant type and delayed for three. Straightforward to add for the User case if reviewers want revocation of a named user to take effect immediately.
  • No feature flag. The blast radius is one TTL, the service fails closed, and CACHE_VERSION invalidates everything in one line.

No frontend query requests userPermissions yet: it was split into !240850 (closed) so the backend ships a release ahead, per GraphQL multi-version sequencing. Tracked in #601817.

References

Edited by Chaya Danzinger

Merge request reports

Loading
Loading