Stop deferring mobile push deliveries on database health signals
What does this MR do and why?
This MR removes the defer_on_database_health_signal call from Todos::PushNotificationWorker. The worker no longer pauses when the shared database health-status indicators return a Stop signal. It adds a documented RuboCop exception for Sidekiq/EnforceDatabaseHealthSignalDeferral and a spec asserting defer_on_database_health_signal? is now false. No other files change.
The incident
On 2026-09-10, a to-do (id 761572397, action marked) was created at 14:33:52.523 UTC via "Add a to-do" on #627724. The recipient has a registered iOS device. The APNs push reached the phone at approximately 14:41 UTC, 6 minutes 54 seconds after the to-do was created. Expected latency for this feature is seconds, not minutes.
Root cause
Sidekiq job log (Kibana, index pubsub-sidekiq-inf-gprd-*, query json.class:"Todos::PushNotificationWorker", all rows share jid 8ae491b62b1519e29bcc4fe0), times UTC:
| Time | Status | Deferred by | Deferred count | Queue duration (s) | Notes |
|---|---|---|---|---|---|
| 14:33:52.551 | start | 0.001 | |||
| 14:33:53.466 | deferred | database_health_check | 1 | 0.001 | |
| 14:35:04.083 | start | 1 | 0.004 | ||
| 14:35:05.769 | deferred | database_health_check | 2 | 0.004 | |
| 14:37:00.999 | start | 2 | 0.002 | ||
| 14:37:02.452 | deferred | database_health_check | 3 | 0.002 | |
| 14:40:44.746 | start | 3 | 0.028 | ||
| 14:40:46.684 | done | 3 | 0.028 | results.delivered = 1 |
Queue durations stayed under 30 ms throughout, so Sidekiq capacity was never the bottleneck. All of the delay came from the deferral middleware.
Database health-status log (same index, query json.job_class_name:"Todos::PushNotificationWorker" and json.indicator_signal:"Stop"):
| Time | Indicator | Reason |
|---|---|---|
| 14:33:53.350 | Gitlab::Database::HealthStatus::Indicators::WriteAheadLog |
WAL archive queue is too big |
| 14:35:05.391 | Gitlab::Database::HealthStatus::Indicators::WriteAheadLog |
WAL archive queue is too big |
| 14:37:01.768 | Gitlab::Database::HealthStatus::Indicators::WriteAheadLog |
WAL archive queue is too big |
WriteAheadLog compares pg_current_wal_insert_lsn() against pg_stat_archiver.last_archived_wal on the main database primary and returns Stop when more than 42 WAL segments are pending archival. This is an instance-wide condition of the primary. It has nothing to do with this worker, its tables, or its load, and it paused every low-urgency worker that opts into health deferral at the same time.
The roughly 7-minute delay comes from the incremental backoff in Gitlab::SidekiqMiddleware::SkipJobs: the first deferral re-enqueues after the worker's delay_by (60 seconds), and each further deferral doubles the delay, minus up to 10% jitter, capped at 30 minutes. This behavior is behind the incremental_database_health_defer_delay feature flag, fully enabled on gprd since 2026-08-11 (rollout tracked in #608150). The observed gaps of 71 s, 115 s, and 222 s match 60 s, then 120 s minus jitter, then 240 s minus jitter, plus scheduler polling. Had the WAL backlog lasted 15 minutes instead of 3, the push would have arrived more than 30 minutes late.
Why remove rather than tune
!254219 (merged) (merged 2026-09-09, in production since 2026-09-09 17:52 UTC) already narrowed the deferred table list from [:todos, :mobile_device_push_subscriptions] to [:mobile_device_push_subscriptions], after deliveries on 2026-09-08 were held back in up-to-30-minute hops by autovacuum on the large todos table. That MR deliberately kept deferral active for the database-wide indicators as an incident lever. Within one day of reaching production, one of those indicators fired and reproduced the same class of delay. Narrowing the table list fixed one trigger; it left the mechanism, and the three other indicators, in place.
Tuning further at the level of a single worker is not possible:
- Indicators cannot be selected per worker.
Gitlab::Database::HealthStatus.evaluatealways runs all four default indicators (AutovacuumActiveOnTable,WriteAheadLog,PatroniApdex,WalRate).defer_on_database_health_signalonly parametrizes schema, tables (used by the autovacuum indicator only), and delay. Overridingdefer_on_database_health_signal?to returnfalseis equivalent to removing the call; changing the shared middleware to add per-indicator opt-out for one worker would be out of proportion to the problem. - The job's database footprint is negligible. It runs
Todo.id_in(todo_ids).pending.with_preloaded_user_and_push_subscriptions, a primary-key lookup of a handful of rows, read from a replica because the worker usesdata_consistency :delayed. The only write issubscription.destroywhen APNs reports a dead device token (:bad_token). The per-user rate limit lives in Redis. The real work is an HTTP/2 request to APNs (worker_has_external_dependencies!). Deferring this job cannot measurably help a primary whose WAL archiver is behind. - Volume is tiny and already gated at the source.
TodoService#enqueue_push_notificationsonly enqueues a job when at least one recipient has a registered device (push_subscribed_todo_ids), and the feature sits behindmobile_push_notifications_dispatch(instance-wide) andmobile_push_notifications(per recipient). Today it is enabled for a handful of dogfooding users. - Deferring defeats the feature. A to-do push is only useful within seconds. Arriving 7 or 30 minutes later means the user has already seen the item on the web, so the notification becomes noise. The health-signal mechanism is designed for heavy background work (its own docs compare it to batched-migration throttling), not for a latency-sensitive notification path.
urgency :highis not an alternative. It is disallowed together withworker_has_external_dependencies!, and APNs latency would not meet the high-urgency p99 target.
Remaining levers
If this worker needs to be held back during an incident, we still have: run_sidekiq_jobs_Todos::PushNotificationWorker (auto-generated feature flag; disabling it defers every job by 5 minutes), drop_sidekiq_jobs_Todos::PushNotificationWorker (drops jobs outright), mobile_push_notifications_dispatch (stops the enqueue at the source), and mobile_push_notifications (per recipient). These give the same incident-response capability without routing through the shared database-health middleware.
Precedent
Disabling Sidekiq/EnforceDatabaseHealthSignalDeferral with a documented reason is established practice elsewhere in the codebase: app/workers/reactive_caching/low_urgency_worker.rb, app/workers/merge_requests/refresh/web_hooks_worker.rb (see #571999), ee/app/workers/search/expire_finder_cache_worker.rb ("Worker only writes to Redis"), and app/workers/database/background_operation/cron_enqueue_worker.rb. 188 low-urgency workers currently run without deferral, listed in .rubocop_todo/sidekiq/enforce_database_health_signal_deferral.yml. Background on the cop: https://docs.gitlab.com/development/sidekiq/#deferring-sidekiq-workers
Risk
Database risk: none measurable, per points 2 and 3 above. Product risk: none; behavior only changes while a health Stop signal is active, when the job now runs immediately instead of sleeping. Rollback is a straight revert of this MR.
This needs sign-off from a maintainer familiar with the Sidekiq health-deferral middleware; group::database health owns incremental_database_health_defer_delay. Please push back if any of this reasoning has a gap.
Verification after deploy
- Thanos:
sum by (reason) (sidekiq_jobs_skipped_total{env="gprd",worker="Todos::PushNotificationWorker",action="deferred"})should stop increasing. - Thanos:
mobile_push_notification_delivery_seconds(buckets 1, 5, 15, 60, 300, 900 s) should land in the 1-5 s buckets. Today's delivery landed in the 900 s bucket. - Kibana: the same queries used above should show no
job_status: deferredrows for this worker going forward.
Kibana and Thanos links below require GitLab team Okta access; screenshots are included in this description for everyone else.
- Kibana, Sidekiq job rows: https://log.gprd.gitlab.net/app/discover#/?_g=(time:(from:'2026-09-10T14:30:00.000Z',to:'2026-09-10T14:50:00.000Z'))&_a=(index:'AWNABDRwNDuQHTm2tH6l',columns:!(json.jid,json.job_status,json.job_deferred_by,json.deferred_count,json.queue_duration_s,json.extra.todos_push_notification_worker.results),query:(language:kuery,query:'json.class:"Todos::PushNotificationWorker"'),sort:!(!(json.time,asc)))
- Kibana, health-status rows: https://log.gprd.gitlab.net/app/discover#/?_g=(time:(from:'2026-09-10T14:30:00.000Z',to:'2026-09-10T14:50:00.000Z'))&_a=(index:'AWNABDRwNDuQHTm2tH6l',columns:!(json.job_class_name,json.health_status_indicator,json.signal_reason,json.status_checker_id),query:(language:kuery,query:'json.job_class_name:"Todos::PushNotificationWorker" and json.indicator_signal:"Stop"'),sort:!(!(json.time,asc)))
- Thanos: https://thanos.gitlab.net/graph?g0.expr=sum%20by%20(reason)%20(sidekiq_jobs_skipped_total%7Benv%3D%22gprd%22%2Cworker%3D%22Todos%3A%3APushNotificationWorker%22%2Caction%3D%22deferred%22%7D)&g0.tab=0&g0.range_input=6h&g0.end_input=2026-09-10%2016%3A00%3A00
References
- !254219 (merged) (previous narrowing of the deferred table list)
- !248026 (merged) (introduced the worker and the dispatch path)
- #607603 (rollout of the
mobile_push_notificationsand dispatch flags) - #607602 (rollout of
mobile_push_registration_api) - #608150 (incremental backoff rollout)
- #536042 (closed) (incremental backoff feature issue)
- https://docs.gitlab.com/development/sidekiq/#deferring-sidekiq-workers (deferral cop documentation)
Screenshots or screen recordings
Kibana and Thanos are only reachable with GitLab team Okta access. The screenshots below are for everyone else.
How to set up and validate locally
bundle exec rspec spec/workers/todos/push_notification_worker_spec.rbbundle exec rubocop app/workers/todos/push_notification_worker.rb spec/workers/todos/push_notification_worker_spec.rb- In
bundle exec rails console:Todos::PushNotificationWorker.defer_on_database_health_signal?returnsfalse(it returnedtruebefore this change). - Optional end to end: enable
mobile_push_notifications_dispatchandmobile_push_notifications, register a device viaPOST /api/v4/user/push_subscriptions, stubGitlab::Database::HealthStatus.evaluateto return a Stop signal, and create a to-do. The job now runs and logsjob_status: doneinstead ofjob_status: deferred.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

