[FF] enable_puma_gvl_metrics -- instrument GVL wait time in Puma

Summary

Roll out GVL wait instrumentation for Puma, currently behind the enable_puma_gvl_metrics feature flag.

When enabled, Puma processes measure Global VM Lock wait. Web and API request logs gain a gvl_thread_wait_s field, and Gitlab::Metrics::Samplers::RubySampler reports ruby_gvl_wait_seconds_total per Puma worker. This is the web counterpart to enable_sidekiq_gvl_metrics, which covers Sidekiq. The two are deliberately separate flags so web and background processing can be rolled out independently.

  • DRI: @hfyngvason
  • Team Slack channel: #g_tenant-services

Note

Process and guidance live in the docs — this issue is just the commands and a place to track the rollout. "Rolling out" means incrementally enabling the flag on GitLab.com to validate stability — it is not the same as releasing the feature, which happens when the flag is removed. Feature flag controls · Feature flag lifecycle

What could go wrong?

Measuring GVL wait installs a Ruby internal thread event hook that fires on every thread context switch. The gem documents a 1-5% slowdown on a fully saturated multi-threaded process, so the risk here is added web request latency rather than data loss or incorrect behaviour. Nothing user-visible changes and no data is written.

The flag is scoped to Feature.current_pod, so a percentage rollout enables it on a percentage of pods rather than fleet-wide. That is the main safety control: latency can be compared between pods that have it on and pods that do not, at the same point in time.

Watch the web and api service overview dashboards on https://dashboards.gitlab.net for request latency and CPU saturation. Roll back if apdex degrades.

A secondary consideration is log volume: gvl_thread_wait_s adds one numeric field per request log line.

Rollout

Run all production /chatops in #production and cross-post the results to #g_tenant-services. Background: incremental rollout process, feature actors.

Non-production

/chatops gitlab run feature set enable_puma_gvl_metrics 50 --actors --dev --pre --staging --staging-ref
/chatops gitlab run feature set enable_puma_gvl_metrics true --dev --pre --staging --staging-ref

Production — percentage rollout (wait ≥15 min between steps, watch dashboards):

/chatops gitlab run feature set enable_puma_gvl_metrics <percentage> --actors

Suggested steps: 1, 10, 25, 50, 100.

Validation

After enabling on any environment, confirm both outputs appear:

  1. The sampler metric reports from Puma workers. RubySampler runs every 60s, so allow a minute:

    sum by (pid) (ruby_gvl_wait_seconds_total{type="web"})

    Expect one series per Puma worker, increasing over time.

  2. Web and API request logs carry the field. In Kibana, filter json.gvl_thread_wait_s: * on the web and api indices.

Note that GVL wait counts thread-seconds rather than elapsed time, so a busy multi-threaded process accumulates it faster than real time. Compare against the thread count of the process, not against wall clock.

Before global rollout

Confirm the relevant gotchas before going to 100% — see enabling a feature for GitLab.com:

Cleanup

Remove the flag once deemed stable — see cleaning up. Track it here, or open a follow-up Feature Flag Cleanup issue. Remove the flag and its YAML definition from the codebase, then:

/chatops gitlab run release check <merge-request-url> <milestone>
/chatops gitlab run feature delete enable_puma_gvl_metrics --dev --pre --staging --staging-ref --production

Because the instrumentation carries a measurable overhead, "cleanup" here may mean deciding to keep the flag as a permanent ops control rather than removing it, in the same way enable_sidekiq_gvl_metrics is treated.

Rollback

/chatops gitlab run feature set enable_puma_gvl_metrics false                                         # production
/chatops gitlab run feature set enable_puma_gvl_metrics false --dev --pre --staging --staging-ref     # non-production
/chatops gitlab run feature delete enable_puma_gvl_metrics --dev --pre --staging --staging-ref --production  # remove entirely

References