WebHooks::LogExecutionWorker parks DB connections while waiting on per-hook lease
Important
Correction (2026-10-01): the original root cause below (DB connections held across the per-hook exclusive lease) is wrong. The leases do not span transactions. The real cause is CPU-bound, payload-size-scaling work in the WebHookLog write. See this comment and the fix in !259143 (merged).
Everyone can contribute. Help move this issue forward while earning points, leveling up and collecting rewards.
Summary
On busy hooks, large numbers of WebHooks::LogExecutionWorker jobs each hold a checked-out primary database connection while doing nothing but sleeping to acquire a per-hook Redis lease. This parks connections on the primary (visible as idle or "idle in transaction" in pg_stat_activity) and risks exhausting the connection pool and saturating Sidekiq.
Evidence
Dashboard: idle-in-transaction connections by worker. WebHooks::LogExecutionWorker with state="idle in transaction" stands out, reaching 23 connections around 2026-10-01 15:49 and sustaining 20+ across the window.
Mechanism
- Every webhook delivery enqueues one
WebHooks::LogExecutionWorker(app/services/web_hook_service.rb:199-208). - The worker is
data_consistency :delayed, so it reads from a replica until its first write (app/workers/web_hooks/log_execution_worker.rb:7). LogExecutionService#executerunslog_executionfirst, which callsWebHookLog.create!(app/services/web_hooks/log_execution_service.rb:19-31). That INSERT is a write, so the database load balancer pins the Sidekiq thread's connection to the primary and keeps it sticky for the rest of the job.update_hook_failure_statethen entersin_lockon a per-hook Redis lease namedweb_hooks:update_hook_failure_state:<hook.id>(app/services/web_hooks/log_execution_service.rb:48-62, lock name atapp/services/web_hooks/log_execution_service.rb:76-78).- The lease settings are
LOCK_TTL = 5.seconds,LOCK_SLEEP = 0.25.seconds, andLOCK_RETRY = 25(app/services/web_hooks/log_execution_service.rb:7-9). A job can sleep up to about 6 seconds waiting for the lease. - All concurrent jobs for one hot hook serialize on that single lease. Most of them spend seconds sleeping while still holding the sticky primary connection.
The docstring of in_lock itself warns that it "could potentially eat up all connection pools" (lib/gitlab/exclusive_lease_helpers.rb, around line 27).
Prior work / context
- Reduced
LOCK_TTLfrom 15s to 5s so jobs finish faster (commitbb0c12a2ae980). - Made a failed lease a no-op (skip the status update) instead of raising (commit
bae2b45d454c3). - Added a
max_concurrency_limit_percentage 0.53throttle to the worker. - Added, then reverted, an experiment to hard-cap concurrency behind a feature flag (commits
b0d686a5da965and4b3f2f4cbf0b7). - Related tracking:
Proposed direction
The root cause is holding a database connection while waiting on the lease. In rough priority order:
- Acquire the lease non-blocking (0 retries or a very small sleep) and skip the failure-state update when it is not acquired. The existing rescue already treats a missed lease as acceptable, since another concurrent job will win it.
- Alternatively, do not pin the database connection during the lease wait, for example by moving the
WebHookLog.create!insert to after the lock section so the thread is not stuck on the primary while sleeping.
Verification needed
Confirm against production pg_stat_activity (connections by application_name or worker) that these connections are LogExecutionWorker jobs parked during the lease wait. Check the related capacity-planning issues above before implementing.
