Zoekt observability: the ORR alert rules and their promtool tests

🤖 AI-authored change.

Adds the 8 Zoekt alerts the Operational Readiness Review asks for, and replaces the two hand-written zoekt.yml files with one Jsonnet definition fanned out to both tenants. They cover search apdex and errors, the indexing task queue, node availability and node storage.

The rules live once in libsonnet/alerts/zoekt-alerts.libsonnet, fanned out per tenant by mimir-rules-jsonnet/zoekt-alerts.jsonnet; the autogenerated-*.yml files are its output, committed because ensure-generated-content-up-to-date requires it.

No threshold is hand-picked: the apdex target from ZOEKT_TARGET_S, the storage markers from WATERMARK_LIMIT_*, the drain deadline from APDEX_THRESHOLD_S, and the burn rates and volume gates from the framework's own generated SLO alerts.

Three decisions worth a reviewer's attention

The node-offline rules use a majority quorum, because the exporter runs on every patroni host reading the same zoekt_nodes table — so disagreement between targets is replication lag, not twelve opinions.

ZoektSearchErrorRateHigh is scoped, because gitlab_sli_global_search_total only advances when record_error_rate is called and the zoekt web path never calls it on success, leaving populations that can never record one.

ZoektTaskQueueNotDraining multiplies rather than divides, so it survives a node that stops completing tasks — mechanism in the collapsed section below.

On-call routing is not confirmed with SRE. Labels are team: global_search, as on the existing Zoekt alerts; nothing sets pager: pagerduty.

Reviewer focus: whether ZoektSearchErrorRateHigh should be a scoped raw-counter burn rate at all, or whether global_search should gain errorRateKind so the framework generates the aggregations — and whether the two storage alerts ought to page.

Review order

Merge request 2 of 3 splitting !11489 (closed), declined as too large to review. Each targets the one above: the dashboard onto master, these alerts onto jmason/zoekt-obs-1-dashboard, then 9 prose-only docs/zoekt/ pages.

Mechanism, verification, and the known gaps

The quorum, measured. Over 30 days on gprd an ANY semantic fires 72 node-steps past for: against a majority's 2 — 68 of those 72 were one lagging replica reporting all 36 nodes offline at once. ALL and MAJORITY are indistinguishable on that window once for: is applied, so MAJORITY is chosen for the fan-out result and because it degrades gracefully; the choice between them stays open for Global Search.

The scoped population, measured. The never-successful populations sit at a constant 100% ratio and were 71.9% of the alert's 7-day numerator. Unscoped it would have paged ~40x/week while the healthy population ran at 0.35x its declared budget; scoping alone takes that to 9 pages/week and scoping plus the burn-rate form to 0.

Why the drain rule multiplies. backlog > deadline * rate equals backlog / rate > deadline for a positive rate, and is total where the quotient is not. gitlab_sli_search_zoekt_tasks_apdex_total is incremented only on completion, so a node that stops completing tasks stops producing the denominator: under a dividing form the join drops it and a firing alert resolves with the backlog untouched. The or on(node_id) (... * 0) arm supplies an explicit zero rate, which reduces the predicate to backlog > 0 — the correct answer for a backlog that never drains — and makes $value the backlog in tasks rather than a duration that could render +Inf.

Each alert is asserted both to fire under its intended condition and to stay silent under the adjacent one; the quorum, both burn-rate window pairs, both volume gates and the drain deadline are pinned by value, and the drain rule's three no-data paths each have a block. The generated YAML was regenerated here rather than copied, promtool and scripts/validate-alerts both pass, and every threshold's measurement — the per-window ratios, the volume-gate rates and the 30-day quorum comparison — is in this snippet, measured through promtool rather than computed by hand.

What I did not verify

  • No alert was fired against a synthetic condition in a live Alertmanager; the promtool tests exercise rule and annotation rendering, not live firing.
  • The promtool tests are not wired into CI — the CI image has no promtool. test/mimir-rules/README.md says what that would take and why it belongs in its own MR.
  • test/mimir-rules/README.md also lists the assertions deliberately not made, including the apdex rule's unkillable > 0 guard and the fast volume gate's 1 → 0.5 case.

🤖 Automated change. Mention @johnmason for feedback, or reply #human to escalate to John.

Edited by John Mason

Merge request reports

Loading
Loading