WebHooks::LogExecutionWorker parks DB connections while waiting on per-hook lease

Important

Correction (2026-10-01): the original root cause below (DB connections held across the per-hook exclusive lease) is wrong. The leases do not span transactions. The real cause is CPU-bound, payload-size-scaling work in the WebHookLog write. See this comment and the fix in !259143 (merged).

Everyone can contribute. Help move this issue forward while earning points, leveling up and collecting rewards.

Summary

On busy hooks, large numbers of WebHooks::LogExecutionWorker jobs each hold a checked-out primary database connection while doing nothing but sleeping to acquire a per-hook Redis lease. This parks connections on the primary (visible as idle or "idle in transaction" in pg_stat_activity) and risks exhausting the connection pool and saturating Sidekiq.

Evidence

Dashboard: idle-in-transaction connections by worker. WebHooks::LogExecutionWorker with state="idle in transaction" stands out, reaching 23 connections around 2026-10-01 15:49 and sustaining 20+ across the window.

idle in transaction connections by worker

Mechanism

  1. Every webhook delivery enqueues one WebHooks::LogExecutionWorker (app/services/web_hook_service.rb:199-208).
  2. The worker is data_consistency :delayed, so it reads from a replica until its first write (app/workers/web_hooks/log_execution_worker.rb:7).
  3. LogExecutionService#execute runs log_execution first, which calls WebHookLog.create! (app/services/web_hooks/log_execution_service.rb:19-31). That INSERT is a write, so the database load balancer pins the Sidekiq thread's connection to the primary and keeps it sticky for the rest of the job.
  4. update_hook_failure_state then enters in_lock on a per-hook Redis lease named web_hooks:update_hook_failure_state:<hook.id> (app/services/web_hooks/log_execution_service.rb:48-62, lock name at app/services/web_hooks/log_execution_service.rb:76-78).
  5. The lease settings are LOCK_TTL = 5.seconds, LOCK_SLEEP = 0.25.seconds, and LOCK_RETRY = 25 (app/services/web_hooks/log_execution_service.rb:7-9). A job can sleep up to about 6 seconds waiting for the lease.
  6. All concurrent jobs for one hot hook serialize on that single lease. Most of them spend seconds sleeping while still holding the sticky primary connection.

The docstring of in_lock itself warns that it "could potentially eat up all connection pools" (lib/gitlab/exclusive_lease_helpers.rb, around line 27).

Prior work / context

Proposed direction

The root cause is holding a database connection while waiting on the lease. In rough priority order:

  • Acquire the lease non-blocking (0 retries or a very small sleep) and skip the failure-state update when it is not acquired. The existing rescue already treats a missed lease as acceptable, since another concurrent job will win it.
  • Alternatively, do not pin the database connection during the lease wait, for example by moving the WebHookLog.create! insert to after the lock section so the thread is not stuck on the primary while sleeping.

Verification needed

Confirm against production pg_stat_activity (connections by application_name or worker) that these connections are LogExecutionWorker jobs parked during the lease wait. Check the related capacity-planning issues above before implementing.

Edited by Hordur Freyr Yngvason