Keep mobile push deliveries timely and log APNs rejections

What does this MR do and why?

Two operational problems surfaced during the first production rollout of mobile push notifications on GitLab.com (rollout issues #607602 and #607603), and this MR fixes both plus a documentation gap:

  1. Pushes were held back for tens of minutes by the database health deferral. Todos::PushNotificationWorker opted into defer_on_database_health_signal with todos in its autovacuum table list. todos is one of the largest and most frequently vacuumed tables, so the worker was deferred whenever a vacuum ran there: sidekiq_jobs_skipped_total{worker="Todos::PushNotificationWorker", action="deferred", reason="database_health_check"} climbed from 1 to 13 within twenty minutes for two jobs, and the corresponding notifications did not reach the device while the signal persisted. A push that arrives half an hour late has lost its purpose.

    The worker never writes to todos. It loads a handful of rows by primary key from a replica (data_consistency :delayed) and its only write is deleting a subscription row when APNs reports a dead device token. It therefore cannot compete with a vacuum on todos, and deferring it on that table only costs latency. This MR keeps the deferral (the Sidekiq/EnforceDatabaseHealthSignalDeferral cop requires it for low-urgency workers, and the database-wide WAL and Patroni apdex indicators remain a useful incident lever) but limits the autovacuum table list to mobile_device_push_subscriptions, the one table the worker writes to.

  2. APNs rejections were invisible. Gitlab::MobilePush::ApnsClient#push returns :failed for any non-2xx APNs answer and for a missing response, but only exceptions were reported to Sentry; the HTTP status and Apple's reason were discarded. Diagnosing a 403 BadEnvironmentKeyInToken (a provider key restricted to the other APNs environment) today required reproducing the send outside GitLab. The client now logs every rejection through Gitlab::AppLogger with apns_status, apns_reason, apns_environment, subscription_id and the resulting outcome: warn for failures, info for the expected dead-token evictions, and a no response reason when APNs does not answer. Device tokens are never logged.

  3. Docs: doc/api/mobile_push_subscriptions.md only named the registration flag. It now also documents that delivery is gated by mobile_push_notifications_dispatch (instance-wide) and mobile_push_notifications (per user), matching the charts and Omnibus documentation.

Impact

  • No change to how to-dos are created, listed, counted or resolved. TodoService is untouched; the enqueue seam it already carries (one indexed query on mobile_device_push_subscriptions after commit, one job only when a recipient has a registered device) is unchanged.
  • The worker's database footprint is unchanged: primary-key reads on todos from a replica, an occasional single-row delete on mobile_device_push_subscriptions. During a database incident it keeps deferring on the WAL and apdex indicators as before; it only stops deferring on todos autovacuum.
  • One additional structured log line per rejected APNs send. Volume is bounded by rejections.

How to set up and validate locally

  1. Run the specs: bin/rspec spec/lib/gitlab/mobile_push/apns_client_spec.rb spec/workers/todos/push_notification_worker_spec.rb
  2. With APNs credentials configured in gitlab.yml (mobile_push.apns), register a device and create a to-do while the three flags are enabled; a rejected send now appears in log/application_json.log as APNs rejected mobile push notification with apns_status and apns_reason.

MR acceptance checklist

This checklist encourages us to confirm any changes have been analyzed to reduce risks in quality, performance, reliability, security, and maintainability.

Merge request reports

Loading
Loading