Bound the checkpoint blob fetch at the current_thread group

What does this MR do and why?

This MR bounds the channel_values blob fetch to the current compaction group, instead of the full ancestor chain. In Kibana on 2026-09-07, p50 db_duration_s for duoWorkflowCheckpoint is about 0.57s with dw_read_blobs_graphql on, versus about 0.08s off. That gap grows with conversation length.

Before and after

Before: every read fetched the blobs of the entire ancestor chain. #full_ancestor_thread_ts walks from the target checkpoint back to the root. #accumulated_blobs_for loads every blob on that path — the rows get fetched, their zlib payloads get detoasted, and the result ships to Rails. Only then does ChannelValuesReconstructor#fold discard, per channel, everything before that channel's newest compaction snapshot. A chat with about 200 checkpoints and three compactions pulls roughly 200 checkpoints of blobs to keep about one group's worth. The waste grows with chat length.

After: the chain walk stops early. Each header carries current_thread, which the gateway bumps only at group starts (compaction or resume). Every group start writes a full snapshot of every channel. So the contiguous run of ancestors sharing the target's counter is exactly the current group, and its blobs alone rebuild identical state.

Concrete example. Chain, newest first: c9(t=2) c8(t=2) c7(t=2) | c6(t=1) … c1(t=0), reading c9:

  • Before: query thread_ts IN (c1..c9). The fold discards c1–c6 per channel in Ruby, after fetch and detoast.
  • After: take_while keeps c9,c8,c7. Query thread_ts IN (c7,c8,c9). Fold output is identical, because c7 (the group start) holds every channel's snapshot.

The SQL shape and index stay the same. Only the IN-list gets shorter, so Postgres reads and detoasts fewer rows.

Correctness guard

The bound relies on the writer contract, so each read checks it: every channel_keys channel must have a compaction anchor in the bounded set. This became viable after ai-assist commit 2219e36b2 (2026-08-14), which made the gateway blob every channel; channel_keys now mirrors channel_values.

On a miss — broken counter numbering, or a group started by a pre-2026-08-14 gateway — the code refetches the full chain (old behavior) and logs a structured warning: "Duo Workflow blob group missing a compaction anchor; refetching the full chain". Wrong data costs one retry, never a wrong fold. Headers without channel_keys skip the bound entirely and keep the unbounded read. Pre-08-14 rows age out of the 30-day retention around 2026-09-13; until then, the log line quantifies how often the fallback fires.

What does not change

  • The write path and schema.
  • Trace and history folds (channel_history, channel_changes), which need the full chain by design.
  • #reconstructed_channel, used for single scalars like status.

Both the single-checkpoint read and the batched page read (blobs_by_thread_ts_for) get the bound. A batched-path guard violation refetches only the offending checkpoint, not the whole page.

How to verify

New specs in ee/spec/models/ai/duo_workflows/workflow_spec.rb cover the bound and the guard, plus pre-existing specs now exercise the fallback path.

  • Bounded query excludes prior-group thread_ts values (checked with QueryRecorder).
  • Bounded fold equals the unbounded fold for mid-stream compactions, scalar rewrites, cross-group resumes, and manual-retry forks.
  • Guard refetch and logging are covered, including a scalar-only bounded set and a membership-less header.
  • Pre-existing #619496 (closed) stale-counter-resume specs now exercise the guard fallback and pass unchanged.

Full suites are green: workflow model spec (527 examples), workflows_internal plus presenter (145), workflows REST (468).

Closes #627866 (closed)

Acceptance criteria

  • Bounded walk implemented for the channel_values path; fold output matches the unbounded read (specs compare bounded vs. unbounded across the chain shapes above).
  • Invariant guard with unbounded refetch and log line, covered by specs.
  • Kibana p50/p95 db_duration_s and cpu_s re-measured after rollout.
Edited by Eduardo Bonet

Merge request reports

Loading
Loading