Zoekt observability: the ORR dashboard and its reducer guards
Adds the Zoekt observability dashboard the Operational Readiness Review asks for, plus the two guards that keep its panel reducers correct. zoekt-main already covers the service-catalog SLIs and container health; the search, indexing, node-capacity and index-state signals had no dashboard.
| File | What it is |
|---|---|
dashboards/zoekt/observability.dashboard.jsonnet |
New zoekt-observability dashboard — 5 rows, 18 panels |
dashboards/zoekt/observability.dashboard_test.jsonnet |
Reads the rendered dashboard and pins the reducer on all 13 exporter-sourced targets |
test/mimir-rules/zoekt-dashboard-reducers_test.yml |
promtool expression tests pinning why max is right |
The search_zoekt_* gauges are reported once per patroni host, so the panels reading them use max; Rails-emitted metrics are not duplicated and stay on sum. Both halves are enforced against the rendered dashboard, so a reducer swap or a dropped grouping label fails make test-jsonnet.
Panel descriptions say what each line means and when it matters.
Known limits
- Three panels are empty on gprd and gstg today. Both tenants pin gitlab-exporter 16.8.0 and the metrics arrived in 16.9.0. The panel titles say they are not reporting yet.
- Index freshness is not instrumented — nothing measures the lag between a write and that write becoming searchable. Row 5 documents the gap rather than shipping an empty panel.
Reviewer focus: whether the panel set is right for an on-caller, and whether the reducer guard belongs in dashboards/ or test/.
Review order
Merge request 1 of 3 splitting !11489 (closed), declined as too large to review. Each targets the one above.
| Order | Merge request | Target | Contents |
|---|---|---|---|
| 1 | this one | master |
Dashboard and its two reducer guards |
| 2 | Zoekt observability: alert rules | jmason/zoekt-obs-1-dashboard |
8 alerts, jsonnet + generated YAML + promtool tests |
| 3 | Zoekt observability: runbooks | jmason/zoekt-obs-2-alerts |
9 docs/zoekt/ pages, prose only |
Mechanism, and verification
Why max. The gitlab-monitor database exporter runs on every patroni host, all querying the same Rails database, so each value is reported once per host — 12 copies on gprd. sum therefore returns ~12x the truth, and sum by (node_name) does not fix it because each node still appears once per fqdn. The two per-row storage panels deduplicate first and then sum; a bare max by (node) there returns ~9% of the node total.
Why two guard files. Neither alone is enough: the promtool file pins the semantics — that sum is 12x, that avg/min differ again, and that a bare max under-reports a node total — but its expressions are string literals, so mutating the dashboard leaves it green. The jsonnet file reads the rendered dashboard, so it catches the mutation but says nothing about why.
Where each drawn threshold comes from.
| Marker | Source |
|---|---|
| Apdex target 15.52s | ZOEKT_TARGET_S, Gitlab::Metrics::GlobalSearchSlis |
| Apdex SLO 0.999, error budget 0.01% | the zoekt service's declared monitoringThresholds |
| Storage watermarks 0.60 / 0.75 / 0.85 | WATERMARK_LIMIT_LOW/HIGH/CRITICAL on Search::Zoekt::Node, production branch — development raises the upper two |
| Indexing apdex target 2h | APDEX_THRESHOLD_S, Gitlab::Metrics::ZoektTasksSlis |
scripts/jsonnet_test.sh dashboards/zoekt/observability.dashboard_test.jsonnet—Passed 12 test cases.promtool test rules test/mimir-rules/zoekt-dashboard-reducers_test.yml—SUCCESS.
What I did not verify
- The dashboard has never been rendered in Grafana;
test-dashboard.shneeds aGRAFANA_API_TOKENand snapshots only work for dashboards already installed.
@johnmason for feedback, or reply #human to escalate to John.