[FF] enable_puma_gvl_metrics -- instrument GVL wait time in Puma
Summary
Roll out GVL wait instrumentation for Puma, currently behind the enable_puma_gvl_metrics feature flag.
When enabled, Puma processes measure Global VM Lock wait. Web and API request logs gain a gvl_thread_wait_s field, and Gitlab::Metrics::Samplers::RubySampler reports ruby_gvl_wait_seconds_total per Puma worker. This is the web counterpart to enable_sidekiq_gvl_metrics, which covers Sidekiq. The two are deliberately separate flags so web and background processing can be rolled out independently.
- DRI: @hfyngvason
- Team Slack channel:
#g_tenant-services
Note
Process and guidance live in the docs — this issue is just the commands and a place to track the rollout. "Rolling out" means incrementally enabling the flag on GitLab.com to validate stability — it is not the same as releasing the feature, which happens when the flag is removed. Feature flag controls · Feature flag lifecycle
What could go wrong?
Measuring GVL wait installs a Ruby internal thread event hook that fires on every thread context switch. The gem documents a 1-5% slowdown on a fully saturated multi-threaded process, so the risk here is added web request latency rather than data loss or incorrect behaviour. Nothing user-visible changes and no data is written.
The flag is scoped to Feature.current_pod, so a percentage rollout enables it on a percentage of pods rather than fleet-wide. That is the main safety control: latency can be compared between pods that have it on and pods that do not, at the same point in time.
Watch the web and api service overview dashboards on https://dashboards.gitlab.net for request latency and CPU saturation. Roll back if apdex degrades.
A secondary consideration is log volume: gvl_thread_wait_s adds one numeric field per request log line.
Rollout
Run all production /chatops in #production and cross-post the results to #g_tenant-services. Background: incremental rollout process, feature actors.
Non-production
/chatops gitlab run feature set enable_puma_gvl_metrics 50 --actors --dev --pre --staging --staging-ref
/chatops gitlab run feature set enable_puma_gvl_metrics true --dev --pre --staging --staging-refProduction — percentage rollout (wait ≥15 min between steps, watch dashboards):
/chatops gitlab run feature set enable_puma_gvl_metrics <percentage> --actorsSuggested steps: 1, 10, 25, 50, 100.
Validation
After enabling on any environment, confirm both outputs appear:
-
The sampler metric reports from Puma workers.
RubySamplerruns every 60s, so allow a minute:sum by (pid) (ruby_gvl_wait_seconds_total{type="web"})Expect one series per Puma worker, increasing over time.
-
Web and API request logs carry the field. In Kibana, filter
json.gvl_thread_wait_s: *on thewebandapiindices.
Note that GVL wait counts thread-seconds rather than elapsed time, so a busy multi-threaded process accumulates it faster than real time. Compare against the thread count of the process, not against wall clock.
Before global rollout
Confirm the relevant gotchas before going to 100% — see enabling a feature for GitLab.com:
- Docs + version history updated
- Breaking changes announced, if any
- Change management issue opened, if required
- External API consumers handled with a fail-open mechanism, if applicable
Cleanup
Remove the flag once deemed stable — see cleaning up. Track it here, or open a follow-up Feature Flag Cleanup issue. Remove the flag and its YAML definition from the codebase, then:
/chatops gitlab run release check <merge-request-url> <milestone>
/chatops gitlab run feature delete enable_puma_gvl_metrics --dev --pre --staging --staging-ref --productionBecause the instrumentation carries a measurable overhead, "cleanup" here may mean deciding to keep the flag as a permanent ops control rather than removing it, in the same way enable_sidekiq_gvl_metrics is treated.
Rollback
/chatops gitlab run feature set enable_puma_gvl_metrics false # production
/chatops gitlab run feature set enable_puma_gvl_metrics false --dev --pre --staging --staging-ref # non-production
/chatops gitlab run feature delete enable_puma_gvl_metrics --dev --pre --staging --staging-ref --production # remove entirelyReferences
- Introduced by !251880
- Depends on !251874 (merged), which adds
ruby_gvl_wait_seconds_total - Sidekiq equivalent flag and its rollout: #577455