Stop deferring mobile push deliveries on database health signals

What does this MR do and why?

This MR removes the defer_on_database_health_signal call from Todos::PushNotificationWorker. The worker no longer pauses when the shared database health-status indicators return a Stop signal. It adds a documented RuboCop exception for Sidekiq/EnforceDatabaseHealthSignalDeferral and a spec asserting defer_on_database_health_signal? is now false. No other files change.

The incident

On 2026-09-10, a to-do (id 761572397, action marked) was created at 14:33:52.523 UTC via "Add a to-do" on #627724. The recipient has a registered iOS device. The APNs push reached the phone at approximately 14:41 UTC, 6 minutes 54 seconds after the to-do was created. Expected latency for this feature is seconds, not minutes.

Root cause

Sidekiq job log (Kibana, index pubsub-sidekiq-inf-gprd-*, query json.class:"Todos::PushNotificationWorker", all rows share jid 8ae491b62b1519e29bcc4fe0), times UTC:

Time Status Deferred by Deferred count Queue duration (s) Notes
14:33:52.551 start 0.001
14:33:53.466 deferred database_health_check 1 0.001
14:35:04.083 start 1 0.004
14:35:05.769 deferred database_health_check 2 0.004
14:37:00.999 start 2 0.002
14:37:02.452 deferred database_health_check 3 0.002
14:40:44.746 start 3 0.028
14:40:46.684 done 3 0.028 results.delivered = 1

Queue durations stayed under 30 ms throughout, so Sidekiq capacity was never the bottleneck. All of the delay came from the deferral middleware.

Database health-status log (same index, query json.job_class_name:"Todos::PushNotificationWorker" and json.indicator_signal:"Stop"):

Time Indicator Reason
14:33:53.350 Gitlab::Database::HealthStatus::Indicators::WriteAheadLog WAL archive queue is too big
14:35:05.391 Gitlab::Database::HealthStatus::Indicators::WriteAheadLog WAL archive queue is too big
14:37:01.768 Gitlab::Database::HealthStatus::Indicators::WriteAheadLog WAL archive queue is too big

WriteAheadLog compares pg_current_wal_insert_lsn() against pg_stat_archiver.last_archived_wal on the main database primary and returns Stop when more than 42 WAL segments are pending archival. This is an instance-wide condition of the primary. It has nothing to do with this worker, its tables, or its load, and it paused every low-urgency worker that opts into health deferral at the same time.

The roughly 7-minute delay comes from the incremental backoff in Gitlab::SidekiqMiddleware::SkipJobs: the first deferral re-enqueues after the worker's delay_by (60 seconds), and each further deferral doubles the delay, minus up to 10% jitter, capped at 30 minutes. This behavior is behind the incremental_database_health_defer_delay feature flag, fully enabled on gprd since 2026-08-11 (rollout tracked in #608150). The observed gaps of 71 s, 115 s, and 222 s match 60 s, then 120 s minus jitter, then 240 s minus jitter, plus scheduler polling. Had the WAL backlog lasted 15 minutes instead of 3, the push would have arrived more than 30 minutes late.

Why remove rather than tune

!254219 (merged) (merged 2026-09-09, in production since 2026-09-09 17:52 UTC) already narrowed the deferred table list from [:todos, :mobile_device_push_subscriptions] to [:mobile_device_push_subscriptions], after deliveries on 2026-09-08 were held back in up-to-30-minute hops by autovacuum on the large todos table. That MR deliberately kept deferral active for the database-wide indicators as an incident lever. Within one day of reaching production, one of those indicators fired and reproduced the same class of delay. Narrowing the table list fixed one trigger; it left the mechanism, and the three other indicators, in place.

Tuning further at the level of a single worker is not possible:

  1. Indicators cannot be selected per worker. Gitlab::Database::HealthStatus.evaluate always runs all four default indicators (AutovacuumActiveOnTable, WriteAheadLog, PatroniApdex, WalRate). defer_on_database_health_signal only parametrizes schema, tables (used by the autovacuum indicator only), and delay. Overriding defer_on_database_health_signal? to return false is equivalent to removing the call; changing the shared middleware to add per-indicator opt-out for one worker would be out of proportion to the problem.
  2. The job's database footprint is negligible. It runs Todo.id_in(todo_ids).pending.with_preloaded_user_and_push_subscriptions, a primary-key lookup of a handful of rows, read from a replica because the worker uses data_consistency :delayed. The only write is subscription.destroy when APNs reports a dead device token (:bad_token). The per-user rate limit lives in Redis. The real work is an HTTP/2 request to APNs (worker_has_external_dependencies!). Deferring this job cannot measurably help a primary whose WAL archiver is behind.
  3. Volume is tiny and already gated at the source. TodoService#enqueue_push_notifications only enqueues a job when at least one recipient has a registered device (push_subscribed_todo_ids), and the feature sits behind mobile_push_notifications_dispatch (instance-wide) and mobile_push_notifications (per recipient). Today it is enabled for a handful of dogfooding users.
  4. Deferring defeats the feature. A to-do push is only useful within seconds. Arriving 7 or 30 minutes later means the user has already seen the item on the web, so the notification becomes noise. The health-signal mechanism is designed for heavy background work (its own docs compare it to batched-migration throttling), not for a latency-sensitive notification path.
  5. urgency :high is not an alternative. It is disallowed together with worker_has_external_dependencies!, and APNs latency would not meet the high-urgency p99 target.

Remaining levers

If this worker needs to be held back during an incident, we still have: run_sidekiq_jobs_Todos::PushNotificationWorker (auto-generated feature flag; disabling it defers every job by 5 minutes), drop_sidekiq_jobs_Todos::PushNotificationWorker (drops jobs outright), mobile_push_notifications_dispatch (stops the enqueue at the source), and mobile_push_notifications (per recipient). These give the same incident-response capability without routing through the shared database-health middleware.

Precedent

Disabling Sidekiq/EnforceDatabaseHealthSignalDeferral with a documented reason is established practice elsewhere in the codebase: app/workers/reactive_caching/low_urgency_worker.rb, app/workers/merge_requests/refresh/web_hooks_worker.rb (see #571999), ee/app/workers/search/expire_finder_cache_worker.rb ("Worker only writes to Redis"), and app/workers/database/background_operation/cron_enqueue_worker.rb. 188 low-urgency workers currently run without deferral, listed in .rubocop_todo/sidekiq/enforce_database_health_signal_deferral.yml. Background on the cop: https://docs.gitlab.com/development/sidekiq/#deferring-sidekiq-workers

Risk

Database risk: none measurable, per points 2 and 3 above. Product risk: none; behavior only changes while a health Stop signal is active, when the job now runs immediately instead of sleeping. Rollback is a straight revert of this MR.

This needs sign-off from a maintainer familiar with the Sidekiq health-deferral middleware; group::database health owns incremental_database_health_defer_delay. Please push back if any of this reasoning has a gap.

Verification after deploy

  • Thanos: sum by (reason) (sidekiq_jobs_skipped_total{env="gprd",worker="Todos::PushNotificationWorker",action="deferred"}) should stop increasing.
  • Thanos: mobile_push_notification_delivery_seconds (buckets 1, 5, 15, 60, 300, 900 s) should land in the 1-5 s buckets. Today's delivery landed in the 900 s bucket.
  • Kibana: the same queries used above should show no job_status: deferred rows for this worker going forward.

Kibana and Thanos links below require GitLab team Okta access; screenshots are included in this description for everyone else.

References

Screenshots or screen recordings

Kibana and Thanos are only reachable with GitLab team Okta access. The screenshots below are for everyone else.

Before After
Sidekiq job log Sidekiq log: three deferrals then done
Three deferred rows before the job finally runs and delivers.
Nothing to show until this deploys. After deploy, check that sum by (reason) (sidekiq_jobs_skipped_total{env="gprd",worker="Todos::PushNotificationWorker",action="deferred"}) on Thanos stops increasing.
Health-status log Health-status log: WriteAheadLog, WAL archive queue is too big
Each deferral is stamped with the same WriteAheadLog Stop signal.
Nothing to show until this deploys. After deploy, Kibana should show no further job_status: deferred rows for this worker, and mobile_push_notification_delivery_seconds should move into the 1-5 s buckets.

How to set up and validate locally

  1. bundle exec rspec spec/workers/todos/push_notification_worker_spec.rb
  2. bundle exec rubocop app/workers/todos/push_notification_worker.rb spec/workers/todos/push_notification_worker_spec.rb
  3. In bundle exec rails console: Todos::PushNotificationWorker.defer_on_database_health_signal? returns false (it returned true before this change).
  4. Optional end to end: enable mobile_push_notifications_dispatch and mobile_push_notifications, register a device via POST /api/v4/user/push_subscriptions, stub Gitlab::Database::HealthStatus.evaluate to return a Stop signal, and create a to-do. The job now runs and logs job_status: done instead of job_status: deferred.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

🤖 Generated with Claude Code

Merge request reports

Loading
Loading