Add Patroni leaderless-cluster and replica-streaming alerts (INC-14518)
What
This adds two paging alerts for a Patroni cluster that has lost its leader, and a non-paging guard for when their inputs disappear. All three are generated for patroni, patroni-ci, patroni-sec and patroni-registry, in gprd and gstg.
| Alert | Fires when (for: 2m) |
Source |
|---|---|---|
PatroniNoHealthyPrimary |
No node passes the Consul check for the <type>-master service |
consul_catalog_service_node_healthy, served by the Consul servers |
PatroniNoReplicaStreaming |
Replicas exist and none has a streaming WAL receiver, and the receiver query is shown to work | pg_stat_wal_receiver_status == 2 on the shard="default" replicas |
PatroniLeadershipAlertsNoData (s3, no page, for: 10m) |
The input of either alert above is absent, so it can't fire | absent() over consul_catalog_service_node_healthy and pg_replication_is_replica |
Each alert has a playbook that sends responders to Primary unresponsive or no Leader before any failover.
Why
This is a corrective action for INC-14518 (production#23045). On Sep 24, Consul marked the primary failed at 22:48:45, and every replica detached at 22:49:56. The primary kept committing. The only page was PatroniScrapeFailures (22:50), and the database failure was confirmed at 23:13 (note_3908190901). Neither alert reads a metric from the primary, whose exporters were down. They would have fired at about 22:51 and 22:52.
Design notes
- Service names:
service_nameequals the cluster type (gitlab-patronirecipes/consul.rb, gprd roles). - Replica signal:
2meansstreaming, and a detached replica's series disappears (gitlab-exporterstemplates/postgres_exporter/queries.yaml.erb). for: 2m: rides through a switchover or a successful automatic failover (TTL of up to 90 s plus promotion).- Excluded replicas: PITR, delayed and archive nodes carry their own
type(for examplepostgres-pitr,postgres-delayed,patroni-ci-archive), and backup nodes areshard="backup". Labels in Mimir come from GCE labels (config-mgmt), not from chefprometheus.labels. A node being rebuilt from the archive (nostreamoverlay) can't hide a detachment, because it doesn't stream. - Noise: the replica alert needs every replica to be down, so one replica out for maintenance doesn't page. A missing receiver series can also mean a broken query, so the rule also needs proof the query works: a replica streaming elsewhere in the environment, or this cluster's
maxreplay lag above 30 s (for every cluster losing its leader at once). Only a query broken on every node stays silent. Lag isn't the main signal, because the exporter reports 0 for a replica restarted with nothing replayed (GREATEST(0, NULL)), and archive replay keeps lag low. The primary alert needs every consul_exporter pod to agree. - Known limitation, split brain after a failover: if a replica is promoted while the old primary is half-dead, both alerts clear, and
PostgresSplitBraincan't see the old primary while its exporter is down. The playbooks say this and require fencing. A detector that survives a failover is proposed as a follow-up. - Staging routing: gstg alerts end in the
blackholeroute, so this MR adds a narrow route that sends both alerts for gstg to#alerts-nonprod. Staging firings can then be reviewed before tuning. gprd routing is unchanged: incident.io, the database automation channel and#production. - During the recovery runbook:
/masterreturns 200 only on the node that holds the leader lock, paused or not (Patroni 3.3.4api.py).PatroniNoHealthyPrimarytherefore stays on while the cluster is paused with no Leader, and clears once the old primary is Leader again, before resume. Both playbooks say so. - Missing data: only one consul_exporter pod runs. If it's down,
PatroniNoHealthyPrimaryhas no data. The same happens to either alert if a cluster's Consul service ortypelabel is renamed, as at a major-version cutover.PatroniLeadershipAlertsNoDatacatches both.PatroniNoReplicaStreamingdoesn't depend on consul_exporter, so it still covers a leaderless cluster meanwhile. - Out of scope:
patroni-ci-v18, the separate PG18 performance-test cluster (production#22864), has its owntypeand Consul service, so the exact-match selectors exclude it. - Not included: "Patroni paused for too long". It needs Patroni's REST
/metricsscraped. That's tracked in production#23070.
Validation
-
promtool test rules test/mimir-rules/patroni-leadership-alerts_test.yml: 17 cases pass. Each alert fires on the Sep 24 sequence and stays silent for a switchover, for one exporter still seeing the primary, for another cluster failing, for a single replica being out, and for a broken WAL-receiver query, and for a WAL-receiver query broken on every node. It still fires when one replica was restarted (lag 0) and when replicas replay from the archive with low lag; both cases fail against the earlierminlag guard. A streaming backup node doesn't hide a real detachment. The guard fires when the consul_exporter pod goes away, when replicas report under a renamedtype, and once per missing series when both are gone, with amissinglabel naming each. -
mimirtool rules checkpasses on all 8 generated files.alertmanager/test-routing.shpasses (73 cases, including a new gstg route test).validate-alerts,jsonnetfmtandmarkdownlintare clean locally, and CI runs the rest. -
Reviewer check (done by @bshah11, note_3940363034): every series returns data in gprd for all four clusters. A 7-day backtest shows both conditions true only during INC-14518 (paging at about 22:51 and 22:53), with no other occurrences. That backtest used the earlier replica expression. The current one fires in more cases (a restarted replica, archive replay), so it needs a re-run before merge.
Deployment plan
gstg and gprd deploy together. Once merged, the ops.gitlab.net pipeline runs deploy-mimir-rules (mimirtool rules sync for every tenant) and update-alertmanager. gprd pages from merge, provided the backtest re-run on the current expression is also clean, and gstg notifies #alerts-nonprod.
- Before merge: approval from a DBRE (
database_automation), and a heads-up in#g_database_operationsand to the EOC with the three playbook links, so the first page isn't a surprise. - Merge, then confirm the deploy (within 30 min): both jobs succeed on ops.gitlab.net. In Grafana Alerting, group
patroni_leadership_alertsshows inmimir-gitlab-gprdandmimir-gitlab-gstgwith health OK.ALERTS{alertname=~"PatroniNoHealthyPrimary|PatroniNoReplicaStreaming|PatroniLeadershipAlertsNoData"}is empty. - End-to-end check in gstg: during the next planned gstg switchover, nothing should fire. During the staging game day for the recovery runbook, both paging alerts should reach
#alerts-nonprodand clear at "Confirm it's Leader on the same timeline". promtool can't cover the Alertmanager and Slack path, so this does. - Observe for 2 weeks: record every firing in gstg and gprd on production#23045 as true or false positive, with its duration. Then tune
for, the 30 s lag threshold and severity in a follow-up MR. - Rollback: silence the specific alert for immediate relief. To remove the rules, revert this MR:
rules syncdeletes rule groups that are no longer in the files (mimirtool 3.0.4pkg/mimirtool/commands/rules.go:660-670), and the same pipeline restores the previous Alertmanager config.
Done when the deploy is confirmed, the gstg check passes, and 2 weeks pass with no unexplained firing. Item 1 of production#23045 can then close.