Bound the size of a single token in jsonstream
What does this MR do and why?
jsonstream had no bound on the size of a single JSON token. The decoder grows its buffer to fit whatever token it is reading, so one oversized string value in an upstream response was buffered in full inside Workhorse, which serves every request. htmlstream already caps this at 1 MiB, and the reasoning written in its source applies to the JSON path unchanged. This gives jsonstream the same ceiling so the two paths fail the same way.
The mechanism differs because the libraries do. htmlstream uses the HTML tokenizer's own SetMaxBuf; jsontext offers no equivalent, and its decoder discards any error returned alongside bytes it did receive. The bound is therefore enforced by the reader handed to the decoder: each read is clipped to what remains of a 1 MiB allowance for the token in progress, measured as bytes supplied minus the decoder's InputOffset. A token that outgrows the allowance has to request more and is refused with ErrTokenTooLarge.
The limit is per token, not per document. A document far larger than 1 MiB made of ordinary sized tokens streams through exactly as before.
References
- Closes #624126
- Found while validating the
pypi_pep_691_jsonrollout in #592164, though as the issue notes, the same transform already serves npm forwarding in production
Differences
Before: a single 64 MiB string value streamed through Transform completed silently at roughly 400 MiB of peak heap, about 6x amplification, with no error. Only the HTML path refused such input.
After: the same input fails closed with ErrTokenTooLarge at about 5 MiB of peak heap, within a few milliseconds. Large documents made of many small tokens are unaffected: the existing packument test and a new two megabyte, thousands-of-tokens test both pass unchanged.
How to set up and validate locally
cd workhorse && go test ./internal/jsonstream/ -count=1 -vTestTransformBoundsTokenBuffer sends a string value one byte over the limit and expects ErrTokenTooLarge. TestTransformStreamsDocumentsLargerThanTokenBound builds a document past two megabytes out of small tokens and expects it to stream cleanly, which is what pins the bound as per token rather than per document.
To reproduce the memory comparison, stream a 64 MiB string value from a generator rather than a fixture, so the input itself is not held in memory, and read runtime.MemStats.HeapInuse before and after. On this branch it stays around 5 MiB.