Reduce OpenBao requests when loading the secrets page
What does this MR do and why?
Loading /<project>/-/secrets (will) resolve the user's effective OpenBao capabilities twice: once in Projects::SecretsController#check_read_capability! to gate the page, and (eventually when implemented) once in the GraphQL userPermissions resolver. Those are separate requests, possibly on different Puma workers, so strong_memoize and SafeRequestStore can't share the result. Each resolution is three sequential OpenBao round trips: login, capabilities-self, and revoke.
This caches the OpenBao lookup in Rails.cache for 30 seconds, keyed per user and per secrets manager.
Benchmarks
ProjectEffectiveCapabilitiesService#execute on a GDK against a local OpenBao. n=20 per run, same project / secrets manager / user in all three, secrets_manager_paid_experience disabled throughout, cache key deleted before each cold iteration on the new code (measuring the cost of one lookup, independent of who calls it)
| run | min | median | p95 | max |
|---|---|---|---|---|
| cache miss, before | 23.45 | 28.04 | 36.70 | 43.93 |
| cache miss, after | 26.14 | 29.52 | 38.55 | 50.16 |
| cache hit, after | 0.12 | 0.12 | 0.29 | 0.29 |
The miss row is the control: unchanged within run-to-run noise, including the tail, so nothing expensive was added to the path that still calls OpenBao. A hit costs 0.12ms against ~28ms, so ~28ms is saved per lookup the cache absorbs.
OpenBao requests counted from the audit log across page loads:
| OpenBao round trips, one page load | before | after, cold | after, cached |
|---|---|---|---|
| capability lookup (3 round trips each) | 3 | 3 | 0 |
(unchanged) secrets list — LIST detailed-metadata |
2 | 2 | 2 |
| total | 5 | 5 | 2 |
Notes:
▎ Inline-authed requests appear twice in the audit log: the internal authentication, then the operation. So a page load shows 7 log entries for these 5 round trips.
▎ Today only the controller resolves capabilities, so the measured saving is on repeat views. Once the frontend requests userPermissions, a single view will resolve twice; the request spec covers the case inclusive of an assumed FE userPermisisons call, and confirms one lookup instead of two with the cache, which would take a cold cache view from 8 round trips (3 for capability lookup * 2 requests, plus 2 for the secrets list) to 5 (3 for the capability lookup once, plus 2 for the secrets list). The table only shows our results though, which currently does not include the userPermission request.
One capability lookup is three HTTP requests to OpenBao, in sequence:
auth/user_jwt/cel/login— mint a real token (inline auth won't work, since capabilities-self resolves policies by looking the token up in the token store)sys/capabilities-self— the actual questionauth/token/revoke-self— throw the token away
This is why the capability lookup row count is 3 before (would have been 6 if the FE userPermissions call landed before this).
Decisions worth a reviewer's attention
- Entitlement is not cached.
executere-applies the read/write entitlement gates on every call. It changes independently of OpenBao (trial expiry, exhausted credits, lapsed subscription), so caching the combined booleans would freeze a billing decision. - Failures are not cached across requests. The fetch returns
nilon the four OpenBao error classes andskip_nil: truekeeps it out of Redis, so a transient blip can't deny a user for the whole TTL. Within a single request,Gitlab::Cache.fetch_oncepins thenilinSafeRequestStore, so a second caller in the same request does not retry OpenBao unnecessarily. - Staleness is bounded at 30 seconds. A revoked or narrowed grant can keep working that long; Rails can't know when OpenBao policies change, so expiry is the mechanism. The GitLab policy layer is not cached, so a user removed from the project is still denied immediately.
- Key uses
secrets_manager.id. This also drops entries from a manager that was deprovisioned and replaced, and it doesn't assume each resource has only one secrets manager —project_secrets_managersenforces that with a unique index, butgroup_secrets_managershas no unique index ongroup_id.scope_nameis in the key because the two manager tables have colliding IDs;current_user.idis required because the CEL program derives policies from the user. - No explicit invalidation when grants change. A grant attaches to a
User, aRole, aGroup, or aMemberRole. Only the single-user case maps to one cache key — a role or group grant affects every member holding it and can't be enumerated cheaply. Since those three would still depend on expiry, worst-case staleness is 30 seconds either way, so this leaves invalidation out rather than making revocation instant for one grant type and delayed for three. Straightforward to add for theUsercase if reviewers want revocation of a named user to take effect immediately. - No feature flag. The blast radius is one TTL, the service fails closed, and
CACHE_VERSIONinvalidates everything in one line.
No frontend query requests userPermissions yet: it was split into !240850 (closed) so the backend ships a release ahead, per GraphQL multi-version sequencing. Tracked in #601817.