[FF] web_hooks_lock_free_failure_state -- lock-free webhook failure-state updates (no exclusive lease)
Summary
Roll out the feature currently behind the web_hooks_lock_free_failure_state feature flag.
- DRI: @stanhu
- Team Slack channel:
#TODO-dri-team-channel
Note
Process and guidance live in the docs. This issue is just the commands and a place to track the rollout. "Rolling out" means incrementally enabling the flag on GitLab.com to validate stability. It is a separate step from releasing the feature, which happens when the flag is removed. Feature flag controls · Feature flag lifecycle
What could go wrong?
- Blast radius: the flag changes how every ProjectHook and GroupHook updates its auto-disable failure state (
recent_failuresanddisabled_until) after each webhook delivery. It is off by default. When on, the update is a single atomic UPDATE instead of an exclusive Redis lease. - A single webhook that fails at very high concurrency could briefly hot-spot one
web_hooksrow (row-lock waits and dead tuples) while it ramps up to being disabled. This stops once the hook is disabled. - The backoff interval and disable thresholds are now computed in SQL. A mismatch with the old Ruby logic would mis-time auto-disabling.
- Rollback is instant. Turn the flag off to fall back to the lease path. There is no data migration and no data-loss risk.
- Dashboards to watch: the idle-in-transaction connections by worker dashboard should show
WebHooks::LogExecutionWorkeridle-in-transaction connections drop. Also watch the webhooks Sidekiq dashboard and the primary Postgres dashboards.
Events
- Exceptions with web_hooks_lock_free_failure_state:1
- Events with web_hooks_lock_free_failure_state:1
- Error rate and other graphs by modifying the examples in the Visualization Library
Feature Flag events are only logged by default for feature flags marked for the current or future milestones. To enable while the feature flag is active, see https://docs.gitlab.com/development/feature_flags/#logging
Rollout
Run all production /chatops in #production and cross-post the results to #TODO-dri-team-channel. Background: incremental rollout process, feature actors.
Non-production
/chatops gitlab run feature set web_hooks_lock_free_failure_state 50 --actors --dev --pre --staging --staging-ref
/chatops gitlab run feature set web_hooks_lock_free_failure_state true --dev --pre --staging --staging-refProduction: percentage rollout (wait ≥15 min between steps, watch dashboards):
/chatops gitlab run feature set web_hooks_lock_free_failure_state <percentage> --actorsOr target specific actors instead:
/chatops gitlab run feature set --project=gitlab-org/gitlab,gitlab-org/gitlab-foss web_hooks_lock_free_failure_state true
/chatops gitlab run feature set --group=gitlab-org,gitlab-com web_hooks_lock_free_failure_state true
/chatops gitlab run feature set --user=stanhu web_hooks_lock_free_failure_state trueBefore global rollout
Confirm the relevant gotchas before going to 100%. See enabling a feature for GitLab.com:
- Docs + version history updated
- Breaking changes announced, if any
- Change management issue opened, if required
- External API consumers handled with a fail-open mechanism, if applicable
Cleanup
Remove the flag once deemed stable. See cleaning up. Track it here, or open a follow-up Feature Flag Cleanup issue. Remove the flag and its YAML definition from the codebase, then:
/chatops gitlab run release check <merge-request-url> <milestone>
/chatops gitlab run feature delete web_hooks_lock_free_failure_state --dev --pre --staging --staging-ref --productionRollback
/chatops gitlab run feature set web_hooks_lock_free_failure_state false # production
/chatops gitlab run feature set web_hooks_lock_free_failure_state false --dev --pre --staging --staging-ref # non-production
/chatops gitlab run feature delete web_hooks_lock_free_failure_state --dev --pre --staging --staging-ref --production # remove entirely