2.5.0 — the measurement chain lands, and its first job was to correct itself

Rows become an answer. eval/analyze.mjs aggregates per arm and deliberately
stops short of a verdict; eval/verdict.mjs applies six gates whose thresholds
were frozen before the run, including one that reads the operator's journal so a
hand-rejected pass is not silently a zero; eval/report.mjs renders the result as
one self-contained HTML file with no network and no build step. When a gate
withholds the verdict, the arm-comparison chart is not drawn at all — not greyed
out, not captioned — and a test asserts that absence.

The chain's first real finding was about the tool that produced it. `usage`
describes an API RESPONSE, but a response reaches the transcript as several
entries — one per content block — each carrying an identical copy of it. Priced
per entry, the same tokens were charged two to five times over: on 19 frozen
transcripts the run total falls from $3,525.55 to $679.35. The inflation is not
flat across token classes and does not cancel in a ratio of two arms, because
the factor IS the number of content blocks per turn — how many tools the model
reached for, which is the variable the doctrine sets out to move. The
unattributed bucket, which exists to report what falls outside every row, was
itself dropping tool calls and per-request server-tool cost; it no longer does.

So the honest summary of this release is uncomfortable and stays on the record:
EPIC-009's first verdict came out AGAINST the doctrine, it was computed from
figures this release proves were too high, and the skew correlates with the arm.
The chain is ready. The number it produced is not yet an answer.