Add Patroni leaderless-cluster and replica-streaming alerts (INC-14518)

What

This adds two paging alerts for a Patroni cluster that has lost its leader, and a non-paging guard for when their inputs disappear. All three are generated for patroni, patroni-ci, patroni-sec and patroni-registry, in gprd and gstg.

Alert Fires when (for: 2m) Source
PatroniNoHealthyPrimary No node passes the Consul check for the <type>-master service consul_catalog_service_node_healthy, served by the Consul servers
PatroniNoReplicaStreaming Replicas exist and none has a streaming WAL receiver, and the receiver query is shown to work pg_stat_wal_receiver_status == 2 on the shard="default" replicas
PatroniLeadershipAlertsNoData (s3, no page, for: 10m) The input of either alert above is absent, so it can't fire absent() over consul_catalog_service_node_healthy and pg_replication_is_replica

Each alert has a playbook that sends responders to Primary unresponsive or no Leader before any failover.

Why

This is a corrective action for INC-14518 (production#23045). On Sep 24, Consul marked the primary failed at 22:48:45, and every replica detached at 22:49:56. The primary kept committing. The only page was PatroniScrapeFailures (22:50), and the database failure was confirmed at 23:13 (note_3908190901). Neither alert reads a metric from the primary, whose exporters were down. They would have fired at about 22:51 and 22:52.

Design notes

  • Service names: service_name equals the cluster type (gitlab-patroni recipes/consul.rb, gprd roles).
  • Replica signal: 2 means streaming, and a detached replica's series disappears (gitlab-exporters templates/postgres_exporter/queries.yaml.erb).
  • for: 2m: rides through a switchover or a successful automatic failover (TTL of up to 90 s plus promotion).
  • Excluded replicas: PITR, delayed and archive nodes carry their own type (for example postgres-pitr, postgres-delayed, patroni-ci-archive), and backup nodes are shard="backup". Labels in Mimir come from GCE labels (config-mgmt), not from chef prometheus.labels. A node being rebuilt from the archive (nostream overlay) can't hide a detachment, because it doesn't stream.
  • Noise: the replica alert needs every replica to be down, so one replica out for maintenance doesn't page. A missing receiver series can also mean a broken query, so the rule also needs proof the query works: a replica streaming elsewhere in the environment, or this cluster's max replay lag above 30 s (for every cluster losing its leader at once). Only a query broken on every node stays silent. Lag isn't the main signal, because the exporter reports 0 for a replica restarted with nothing replayed (GREATEST(0, NULL)), and archive replay keeps lag low. The primary alert needs every consul_exporter pod to agree.
  • Known limitation, split brain after a failover: if a replica is promoted while the old primary is half-dead, both alerts clear, and PostgresSplitBrain can't see the old primary while its exporter is down. The playbooks say this and require fencing. A detector that survives a failover is proposed as a follow-up.
  • Staging routing: gstg alerts end in the blackhole route, so this MR adds a narrow route that sends both alerts for gstg to #alerts-nonprod. Staging firings can then be reviewed before tuning. gprd routing is unchanged: incident.io, the database automation channel and #production.
  • During the recovery runbook: /master returns 200 only on the node that holds the leader lock, paused or not (Patroni 3.3.4 api.py). PatroniNoHealthyPrimary therefore stays on while the cluster is paused with no Leader, and clears once the old primary is Leader again, before resume. Both playbooks say so.
  • Missing data: only one consul_exporter pod runs. If it's down, PatroniNoHealthyPrimary has no data. The same happens to either alert if a cluster's Consul service or type label is renamed, as at a major-version cutover. PatroniLeadershipAlertsNoData catches both. PatroniNoReplicaStreaming doesn't depend on consul_exporter, so it still covers a leaderless cluster meanwhile.
  • Out of scope: patroni-ci-v18, the separate PG18 performance-test cluster (production#22864), has its own type and Consul service, so the exact-match selectors exclude it.
  • Not included: "Patroni paused for too long". It needs Patroni's REST /metrics scraped. That's tracked in production#23070.

Validation

  • promtool test rules test/mimir-rules/patroni-leadership-alerts_test.yml: 17 cases pass. Each alert fires on the Sep 24 sequence and stays silent for a switchover, for one exporter still seeing the primary, for another cluster failing, for a single replica being out, and for a broken WAL-receiver query, and for a WAL-receiver query broken on every node. It still fires when one replica was restarted (lag 0) and when replicas replay from the archive with low lag; both cases fail against the earlier min lag guard. A streaming backup node doesn't hide a real detachment. The guard fires when the consul_exporter pod goes away, when replicas report under a renamed type, and once per missing series when both are gone, with a missing label naming each.

  • mimirtool rules check passes on all 8 generated files. alertmanager/test-routing.sh passes (73 cases, including a new gstg route test). validate-alerts, jsonnetfmt and markdownlint are clean locally, and CI runs the rest.

  • Reviewer check (done by @bshah11, note_3940363034): every series returns data in gprd for all four clusters. A 7-day backtest shows both conditions true only during INC-14518 (paging at about 22:51 and 22:53), with no other occurrences. That backtest used the earlier replica expression. The current one fires in more cases (a restarted replica, archive replay), so it needs a re-run before merge.

Deployment plan

gstg and gprd deploy together. Once merged, the ops.gitlab.net pipeline runs deploy-mimir-rules (mimirtool rules sync for every tenant) and update-alertmanager. gprd pages from merge, provided the backtest re-run on the current expression is also clean, and gstg notifies #alerts-nonprod.

  1. Before merge: approval from a DBRE (database_automation), and a heads-up in #g_database_operations and to the EOC with the three playbook links, so the first page isn't a surprise.
  2. Merge, then confirm the deploy (within 30 min): both jobs succeed on ops.gitlab.net. In Grafana Alerting, group patroni_leadership_alerts shows in mimir-gitlab-gprd and mimir-gitlab-gstg with health OK. ALERTS{alertname=~"PatroniNoHealthyPrimary|PatroniNoReplicaStreaming|PatroniLeadershipAlertsNoData"} is empty.
  3. End-to-end check in gstg: during the next planned gstg switchover, nothing should fire. During the staging game day for the recovery runbook, both paging alerts should reach #alerts-nonprod and clear at "Confirm it's Leader on the same timeline". promtool can't cover the Alertmanager and Slack path, so this does.
  4. Observe for 2 weeks: record every firing in gstg and gprd on production#23045 as true or false positive, with its duration. Then tune for, the 30 s lag threshold and severity in a follow-up MR.
  5. Rollback: silence the specific alert for immediate relief. To remove the rules, revert this MR: rules sync deletes rule groups that are no longer in the files (mimirtool 3.0.4 pkg/mimirtool/commands/rules.go:660-670), and the same pipeline restores the previous Alertmanager config.

Done when the deploy is confirmed, the gstg check passes, and 2 weeks pass with no unexplained firing. Item 1 of production#23045 can then close.

Edited by Vamshidhar Poralla

Merge request reports

Loading
Loading