Stop reserving 4MB per Duo Workflow WebSocket connection
What does this MR do and why?
newWsManager allocated its protojson marshalling buffer at the full
ActionResponseBodyLimit (4,190,208 bytes) for every WebSocket connection,
even though MarshalAppend only needs capacity and grows on demand. A typical
marshalled action is a few hundred bytes.
The buffer is reachable from the runner for the whole connection, and agent flows stay open for over an hour, so it was never collectable.
The mechanism is GC pacing, not resident pages
This is the part worth reviewing carefully, because the obvious reading is wrong.
make([]byte, 4MB) takes a fresh span that Go never memsets, so the pages are
not dirtied and the allocation is virtual only. Measured directly, 200
connections cost ~0 RSS on their own.
The real cost is that it inflates the live heap. With GOGC=100 and no
GOMEMLIMIT (workhorse sets neither) the GC goal is ~2x live heap, so the
collector effectively stops running and the genuinely dirty transient garbage
accumulates instead of being reclaimed.
Local reproduction, 200 connections plus an identical page-touching workload:
NextGC |
HeapSys |
maxRSS | |
|---|---|---|---|
| eager 4MB buffer | 1653 MB | 1083 MB | 311 MB |
| grown on demand | 53 MB | 75 MB | 102 MB |
Same allocation rate, 3x the RSS. So the win is a multiplier on all transient garbage rather than a fixed 4MB per connection.
The change
- Start the buffer at 4 KiB and let
MarshalAppendgrow it. - Release it after a write that grew it past 256 KiB, so one large action does not pin its capacity for the rest of a multi-hour connection.
BenchmarkNewWsManager: 4,194,594 B/op -> 4,384 B/op (~957x), 145 us -> 430 ns.
Observability
Adds gitlab_workhorse_duo_workflow_connections_open. Concurrency was not
measurable before: connections_total is a counter, and
gitlab_workhorse_http_in_flight_requests is not a usable proxy because it is
an unlabelled Gauge (not a GaugeVec), so it cannot be narrowed to this
route, and it also counts the HTTP actions that re-enter the upstream router
while the WebSocket is still open. Memory per connection cannot be verified
without this gauge.
References
Capacity warning for ai-assisted, kube_container_rss_request:
https://gitlab.com/gitlab-com/gl-infra/capacity-planning-trackers/gitlab-com/-/work_items/2668
The gitlab-workhorse container in the ai-assisted gprd deployment reaches
7.5x its 200M memory request and 75% of its 2G limit. An OOM there kills every
in-flight agent flow WebSocket on the pod.
Screenshots or screen recordings
Not applicable, no UI changes.
How to set up and validate locally
-
Confirm the per-connection allocation dropped:
cd workhorse go test ./internal/ai_assist/duoworkflow -run XXX -bench BenchmarkNewWsManager -benchmemExpect ~4,384 B/op. Before this change it was 4,194,594 B/op.
-
Run the regression tests, which assert the buffer is neither pre-sized to the 4MB ceiling nor left oversized after a large action:
go test ./internal/ai_assist/duoworkflow -run 'TestWsManager_MarshalBufferSizing|TestConnectionsOpen' -v
Verifying in production
Because the mechanism is GC pacing, the primary signal is the GC goal, which the default Go collector already exposes. No new instrumentation is needed for it:
quantile(0.99, go_memstats_next_gc_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})
quantile(0.99, go_memstats_heap_inuse_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})
quantile(0.99, process_resident_memory_bytes{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"})Control for a traffic change rather than a real improvement. The allocation rate should stay flat while RSS falls:
sum(rate(go_memstats_alloc_bytes_total{env="gprd", type="ai-assisted", job=~"gitlab-workhorse.*"}[5m]))
rate(gitlab_workhorse_duo_workflow_connections_total{env="gprd"}[5m])Expected trade-off to watch: a lower heap goal means more frequent collections,
so rate(go_gc_duration_seconds_count[5m]) should rise. Each cycle is much
cheaper (HeapSys 1083 MB -> 75 MB locally), but confirm CPU and
gitlab_workhorse_duo_workflow_http_action_duration_seconds do not regress.
MR acceptance checklist
Evaluate this MR against the MR acceptance checklist. It helps you analyze changes to reduce risks in quality, performance, reliability, security, and maintainability.
Follow-up, deliberately not in this MR
runHTTPActionHandler holds three simultaneous full copies of every action
response body: the bytes.Buffer in nullResponseWriter, the string copy from
nrw.body.String(), and the proto wire encoding. Measured, a 3.28 MB checkpoint
occupies ~10.55 MB of genuinely dirtied heap. That is the memory this MR lets the
GC actually reclaim, but reducing the copies is a separate, larger change.