Stop reserving 4MB per Duo Workflow WebSocket connection

What does this MR do and why?

newWsManager allocated its protojson marshalling buffer at the full ActionResponseBodyLimit (4,190,208 bytes) for every WebSocket connection, even though MarshalAppend only needs capacity and grows on demand. A typical marshalled action is a few hundred bytes.

The buffer is reachable from the runner for the whole connection, and agent flows stay open for over an hour, so it was never collectable.

The mechanism is GC pacing, not resident pages

This is the part worth reviewing carefully, because the obvious reading is wrong. make([]byte, 4MB) takes a fresh span that Go never memsets, so the pages are not dirtied and the allocation is virtual only. Measured directly, 200 connections cost ~0 RSS on their own.

The real cost is that it inflates the live heap. With GOGC=100 and no GOMEMLIMIT (workhorse sets neither) the GC goal is ~2x live heap, so the collector effectively stops running and the genuinely dirty transient garbage accumulates instead of being reclaimed.

Local reproduction, 200 connections plus an identical page-touching workload:

NextGC HeapSys maxRSS
eager 4MB buffer 1653 MB 1083 MB 311 MB
grown on demand 53 MB 75 MB 102 MB

Same allocation rate, 3x the RSS. So the win is a multiplier on all transient garbage rather than a fixed 4MB per connection.

The change

  • Start the buffer at 4 KiB and let MarshalAppend grow it.
  • Release it after a write that grew it past 256 KiB, so one large action does not pin its capacity for the rest of a multi-hour connection.

BenchmarkNewWsManager: 4,194,594 B/op -> 4,384 B/op (~957x), 145 us -> 430 ns.

Observability

Adds gitlab_workhorse_duo_workflow_connections_open. Concurrency was not measurable before: connections_total is a counter, and gitlab_workhorse_http_in_flight_requests is not a usable proxy because it is an unlabelled Gauge (not a GaugeVec), so it cannot be narrowed to this route, and it also counts the HTTP actions that re-enter the upstream router while the WebSocket is still open. Memory per connection cannot be verified without this gauge.

References

Capacity warning for ai-assisted, kube_container_rss_request: https://gitlab.com/gitlab-com/gl-infra/capacity-planning-trackers/gitlab-com/-/work_items/2668

The gitlab-workhorse container in the ai-assisted gprd deployment reaches 7.5x its 200M memory request and 75% of its 2G limit. An OOM there kills every in-flight agent flow WebSocket on the pod.

Screenshots or screen recordings

Not applicable, no UI changes.

How to set up and validate locally

  1. Confirm the per-connection allocation dropped:

    cd workhorse
    go test ./internal/ai_assist/duoworkflow -run XXX -bench BenchmarkNewWsManager -benchmem

    Expect ~4,384 B/op. Before this change it was 4,194,594 B/op.

  2. Run the regression tests, which assert the buffer is neither pre-sized to the 4MB ceiling nor left oversized after a large action:

    go test ./internal/ai_assist/duoworkflow -run 'TestWsManager_MarshalBufferSizing|TestConnectionsOpen' -v

Verifying in production

Because the mechanism is GC pacing, the primary signal is the GC goal, which the default Go collector already exposes. No new instrumentation is needed for it:

quantile(0.99, go_memstats_next_gc_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})
quantile(0.99, go_memstats_heap_inuse_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})
quantile(0.99, process_resident_memory_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})

Control for a traffic change rather than a real improvement. The allocation rate should stay flat while RSS falls:

sum(rate(go_memstats_alloc_bytes_total{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"}[5m]))
rate(gitlab_workhorse_duo_workflow_connections_total{env="gprd"}[5m])

Expected trade-off to watch: a lower heap goal means more frequent collections, so rate(go_gc_duration_seconds_count[5m]) should rise. Each cycle is much cheaper (HeapSys 1083 MB -> 75 MB locally), but confirm CPU and gitlab_workhorse_duo_workflow_http_action_duration_seconds do not regress.

MR acceptance checklist

Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.

Follow-up, deliberately not in this MR

runHTTPActionHandler holds three simultaneous full copies of every action response body: the bytes.Buffer in nullResponseWriter, the string copy from nrw.body.String(), and the proto wire encoding. Measured, a 3.28 MB checkpoint occupies ~10.55 MB of genuinely dirtied heap. That is the memory this MR lets the GC actually reclaim, but reducing the copies is a separate, larger change.

Merge request reports

Loading
Loading