perf(configpublisher): halve HAProxyCfg publication latency

Why republishing the 500 KB HAProxyCfg took 4 s

Measured with apiserver request metrics over TestScale's 20 single-change samples (800 Ingresses, 74 auxiliary children per publication: 61 map files, 10 general files, 1 crt-list, 2 Secrets). Every change creates a new auxiliary set, so all 74 children are recreated, the 74 old ones deleted, and each pod re-stamps all 74. Per change:

cost per change before
full GET haproxycfgs (500 KB each) ~75, ~37 MB transferred and decoded by the leader
child creates / deletes 74 / 74, serial
child status stamps 2 x 74, serial per pod

The 75 full reads came from the cleanup fence: ensurePublicationCurrent re-read the whole HAProxyCfg before every stale-child deletion, plus one read before the spec update and one before the status-reference patch.

What changed (all HAPTIC-side; validation strength unchanged)

  • Cleanup fence reads metadata first. Before each deletion the fence reads the HAProxyCfg's resourceVersion through a metadata client. If it equals the version that last passed the full check, the object is byte-identical to the one that passed, so the verdict is the same; otherwise the full check runs again. Every deletion is still preceded by a passing check, and deletions stay serial.
  • Status-reference patch without a prior read when the references change. The JSON patch's test operations already verify UID, owner references, spec and set ID on the server; the elision path for unchanged references still reads the live object first.
  • Spec update from the informer cache when the cached copy differs from the desired one. The update carries the cached resourceVersion, so a stale copy is refused with a conflict and the retry reads the live object. An unchanged object is still confirmed by a live read.
  • Children and their pod status written in parallel (bound 8, the same as pod cleanup). Names keep input order; a failed write returns only the names that published.

Before / after (TestScale, isolated kind cluster, HAProxy 3.4, same host)

metric before after
HAProxyCfg publication median 4.22 s 2.11 s
HAProxyCfg publication p95 5.19 s 2.68 s
publication from idle (sample 1) 2.06 s 0.65 s
HAProxyCfg bytes read per change ~37 MB ~0.5 MB
create → routed p95 0.23 s 0.28 s (unchanged path, noise)
controller working set 924 MB 933 MB
controller CPU over the run 514 s 506 s
seed → converged 67.4 s 68.3 s

Back-to-back changes still wait on the previous publication's serial cleanup (74 deletions x ~20 ms). The remaining structural cost is that an unchanged auxiliary file is recreated under every new set name; removing that changes the published-set protocol readers verify, and belongs in its own issue.

Verification

  • make test-unit for pkg/k8s/configpublisher: new tests count HAProxyCfg reads per change (2 instead of 10 with 8 children), prove a publication superseded mid-cleanup still stops the next deletion, exercise a fresh and a stale informer copy, and check child order under parallel writes. TestRuntimeConfigStatusConcurrentWriters now injects the concurrent writer at the status patch, the only point left.
  • TestScale before/after as above, exit 0 both.
  • make lint exit 0 and make test exit 0 on the union of this MR and the companion MR (disjoint files), under the shared heavy-gate lock.

Relates to #283 (closed)

Merge request reports

Loading
Loading