Tags give the ability to mark specific points in history as being important
-
v2.5.0
0f3c12db · ·2.5.0 — the measurement chain lands, and its first job was to correct itself Rows become an answer. eval/analyze.mjs aggregates per arm and deliberately stops short of a verdict; eval/verdict.mjs applies six gates whose thresholds were frozen before the run, including one that reads the operator's journal so a hand-rejected pass is not silently a zero; eval/report.mjs renders the result as one self-contained HTML file with no network and no build step. When a gate withholds the verdict, the arm-comparison chart is not drawn at all — not greyed out, not captioned — and a test asserts that absence. The chain's first real finding was about the tool that produced it. `usage` describes an API RESPONSE, but a response reaches the transcript as several entries — one per content block — each carrying an identical copy of it. Priced per entry, the same tokens were charged two to five times over: on 19 frozen transcripts the run total falls from $3,525.55 to $679.35. The inflation is not flat across token classes and does not cancel in a ratio of two arms, because the factor IS the number of content blocks per turn — how many tools the model reached for, which is the variable the doctrine sets out to move. The unattributed bucket, which exists to report what falls outside every row, was itself dropping tool calls and per-request server-tool cost; it no longer does. So the honest summary of this release is uncomfortable and stays on the record: EPIC-009's first verdict came out AGAINST the doctrine, it was computed from figures this release proves were too high, and the skew correlates with the arm. The chain is ready. The number it produced is not yet an answer.
-
-
rejected/task-090-off-1
eaca7d9c · ·Odrzucony przebieg §7: P1 bez podstawionego numeru taska (0XX). Sesja eb076d7c.
-
v2.4.0
dd1129f8 · ·2.4.0 — the measurement can finally have two sides `tune-context experiment` switches the doctrine and RAG off and back on, so the same work can be run in both arms: hook registrations, the managed CLAUDE.md block, doctrine skills, agents and the project .mcp.json. Every touched path is fingerprinted and copied aside before it is modified, so a `set` interrupted halfway is still reversible; `restore` puts the bytes back or refuses, naming what would be lost. `verify` reads the off arm as "off" instead of as a broken install, and `init` refuses while an arm is set. Cost stops being a proxy. eval/prompt-metrics.mjs turns transcripts into one row per prompt — time split into queued and work, tokens split by cache class, model, effort, tools, delegations — and eval/pricing.mjs prices them at API list price, cache multipliers included. It is NOT a bill: a Pro/Max plan charges its flat fee, and the fields say so in their names. The breakdown carries the epic's thesis in one line: on the local baseline, cache reads and cache writes are nearly the whole figure while plain input is 0.07% of it. Rows know which arm they came from. A branch shaped exp/<task>/<arm> says which arm was intended, the TASK-087 ledger which was set, and a disagreement leaves the arm unresolved with the reason on the row rather than picking a winner. pass_order rides along, because the second implementation of a task is cheaper for a reason that has nothing to do with the arm. Pre-release review caught two measurement bugs: `restore` now ends a run even when it parked nothing (the "on" arm's window used to stay open forever, and only that arm's, which made the contamination one-sided), and a cost that is not a finite number is refused instead of being counted as priced and turning the whole run's total into $NaN.
-
v2.3.1
877fbb86 · ·2.3.1 — third context notice at 1.5x, graded by rung Middle rung at 1.5x the threshold (180k by default) joins 120k and 240k, cutting worst-case silence between notices from 120k to 60k. The ladder is multipliers, so TUNE_CONTEXT_REMIND_TOKENS scales all three. Each rung is louder than the last: wait for a natural break, raise it at the next pause anyway, say it this turn. Fixes the second notice calling itself 'second and final', which a three-rung ladder makes false. Docs describing the reminder as firing once (README, tune-context skill) — wrong since 2.1.0, corrected here.
-
v2.3.0
05f48289 · ·v2.3.0 — the adoption report can tell "RAG installed and ignored" from "no RAG here": a project .mcp.json probe behind a soft `~` verdict, guarded by a temporal check (config must predate the sessions), explicit rejections, observation precedence and the PO-frozen session floor now applied per project; search counted per project so one RAG-using repo no longer reads as coverage for the rest
-
v2.2.0
cc572ea6 · ·v2.2.0 — authored CLAUDE.md sections live below the managed block (never regenerated); verify content-compares deployed hooks/agents/skills against the installed package, plus a read-only `tune-context verify` and a guard for files still in the old layout
-
v2.1.0
449d4684 · ·v2.1.0 — context reminder fires a second and final time at 2x the threshold (240k default); systemMessage added best-effort (does not render for UserPromptSubmit)
-
v2.0.1
3d1ae904 · ·v2.0.1 — fix: npm-installed CLI ran nothing (entry-point guard vs extensionless bin symlink)
-
v2.0.0
05d91808 · ·v2.0.0 — install integrity, fact fidelity, and a context reminder that measures EPIC-008: the installer and detector stop producing signal they then measure. Plugin+CLI dual registration is detected and no longer fakes a churn verdict; detect writes the package manager it actually found (pnpm/yarn, through to permissions.allow); CLAUDE_CONFIG_DIR and the allow list have one definition each; rejected .mcp.json servers are honored; PreCompact is fail-safe with collision-free state paths and 30-day retention; cli/ gains unit coverage; docs and CLI UX close the gaps. TASK-079: new context-size reminder (UserPromptSubmit). The model cannot see its own context usage, so a hook reads the transcript tail, sums the tokens actually re-sent, and once past a PO-frozen 120k injects one short note suggesting /compact — once per crossing, re-armed only after a compaction. BREAKING: `tune-context` with no subcommand no longer runs `init`; `init` is now required explicitly. Also adds `tune-context --version`. See CHANGELOG.md, including "Upgrading from 1.1.0".
-
v1.1.0
c8c8c1dc · ·v1.1.0 — CLAUDE_CONFIG_DIR consistency + adoption-report hardening EPIC-006: CLAUDE_CONFIG_DIR honored across the whole hook lifecycle (registration + state paths); verify cross-checks hook wiring; smoke covers both path modes. EPIC-007: adoption report splits code search from shell plumbing, earns its verdict (sample-size + availability aware), and stops printing false negatives. See CHANGELOG.md.
-
-
v1.0.0
d7972c83 · ·tune-context v1.0.0 — first public release (npm + Claude Code plugin marketplace)