Zoekt observability: the ORR dashboard and its reducer guards

🤖 AI-authored change.

Adds the Zoekt observability dashboard the Operational Readiness Review asks for, plus the two guards that keep its panel reducers correct. zoekt-main already covers the service-catalog SLIs and container health; the search, indexing, node-capacity and index-state signals had no dashboard.

File What it is
dashboards/zoekt/observability.dashboard.jsonnet New zoekt-observability dashboard — 5 rows, 18 panels
dashboards/zoekt/observability.dashboard_test.jsonnet Reads the rendered dashboard and pins the reducer on all 13 exporter-sourced targets
test/mimir-rules/zoekt-dashboard-reducers_test.yml promtool expression tests pinning why max is right

The search_zoekt_* gauges are reported once per patroni host, so the panels reading them use max; Rails-emitted metrics are not duplicated and stay on sum. Both halves are enforced against the rendered dashboard, so a reducer swap or a dropped grouping label fails make test-jsonnet.

Panel descriptions say what each line means and when it matters.

Known limits

  • Three panels are empty on gprd and gstg today. Both tenants pin gitlab-exporter 16.8.0 and the metrics arrived in 16.9.0. The panel titles say they are not reporting yet.
  • Index freshness is not instrumented — nothing measures the lag between a write and that write becoming searchable. Row 5 documents the gap rather than shipping an empty panel.

Reviewer focus: whether the panel set is right for an on-caller, and whether the reducer guard belongs in dashboards/ or test/.

Review order

Merge request 1 of 3 splitting !11489 (closed), declined as too large to review. Each targets the one above.

Order Merge request Target Contents
1 this one master Dashboard and its two reducer guards
2 Zoekt observability: alert rules jmason/zoekt-obs-1-dashboard 8 alerts, jsonnet + generated YAML + promtool tests
3 Zoekt observability: runbooks jmason/zoekt-obs-2-alerts 9 docs/zoekt/ pages, prose only
Mechanism, and verification

Why max. The gitlab-monitor database exporter runs on every patroni host, all querying the same Rails database, so each value is reported once per host — 12 copies on gprd. sum therefore returns ~12x the truth, and sum by (node_name) does not fix it because each node still appears once per fqdn. The two per-row storage panels deduplicate first and then sum; a bare max by (node) there returns ~9% of the node total.

Why two guard files. Neither alone is enough: the promtool file pins the semantics — that sum is 12x, that avg/min differ again, and that a bare max under-reports a node total — but its expressions are string literals, so mutating the dashboard leaves it green. The jsonnet file reads the rendered dashboard, so it catches the mutation but says nothing about why.

Where each drawn threshold comes from.

Marker Source
Apdex target 15.52s ZOEKT_TARGET_S, Gitlab::Metrics::GlobalSearchSlis
Apdex SLO 0.999, error budget 0.01% the zoekt service's declared monitoringThresholds
Storage watermarks 0.60 / 0.75 / 0.85 WATERMARK_LIMIT_LOW/HIGH/CRITICAL on Search::Zoekt::Node, production branch — development raises the upper two
Indexing apdex target 2h APDEX_THRESHOLD_S, Gitlab::Metrics::ZoektTasksSlis
  • scripts/jsonnet_test.sh dashboards/zoekt/observability.dashboard_test.jsonnet — Passed 12 test cases.
  • promtool test rules test/mimir-rules/zoekt-dashboard-reducers_test.yml — SUCCESS.

What I did not verify

  • The dashboard has never been rendered in Grafana; test-dashboard.sh needs a GRAFANA_API_TOKEN and snapshots only work for dashboards already installed.

🤖 Automated change. Mention @johnmason for feedback, or reply #human to escalate to John.

Edited by John Mason

Merge request reports

Loading
Loading