Fix channel history fold dropping compaction messages
What
Fixes ChannelValuesReconstructor#channel_history and #channel_changes so they keep
messages added during a compaction-triggering checkpoint step, instead of dropping the
compaction snapshot outright.
A compaction snapshot holds the whole channel value, so when it restates what the fold already holds, the rest is what that step added. Both folds now append that tail. A snapshot that rewrites or replaces the fold instead (a replace-semantics channel, or a trimmed list) says nothing about what's new, so they take none of it.
Workflow#latest_channel_message had the same defect on the preview surface: it
filtered to conversation blobs, and the compaction-triggering step writes none. The
session-list preview sat a turn behind the message list it reads alongside. It now
reads the newest blob whatever its step_action.
Specs cover messages introduced only by the compaction-triggering checkpoint (both folds), a snapshot that rewrites the head (both folds), and the preview keeping up.
Why
A user typed "Say 4", a compaction ran, and the agent replied "4". After a page reload, both messages were gone from the chat. They only came back later, out of order, after the next turn. The cause: those messages lived only inside a compaction snapshot that the fold threw away.
#channel_history feeds GraphQL messages behind dw_read_blobs_graphql;
#channel_changes feeds the trace artifact behind dw_read_blobs_trace. Neither flag
has ever been enabled on GitLab.com. This isn't a production incident; it blocks
turning them on.
Rebased onto the scalar fix
#channel_changes now also carries the scalar-history fix from
#619532 (closed), which landed on master while this MR was open.
The two changes meet in the same method, so the rebase reconciled them by value shape: a list snapshot
contributes its tail (this MR), a dict snapshot is still dropped before decoding, and a scalar records
every transition bar a consecutive repeat (the 619532 fix). Both sets of specs pass together.
Closes #619495 (closed)
Database
No schema change and no migration. One existing query changed: Ai::DuoWorkflows::Workflow#latest_channel_message
dropped its step_action = 'conversation' predicate. Everything else in this MR is in-memory folding.
The predicate was never an index condition. idx_duo_wf_checkpoint_blobs_dedup is
(project_id, workflow_id, thread_ts, channel, version, step_action, workflow_created_at), and the query doesn't
constrain version, so step_action could only ever be a filter. Dropping a conjunct from the filter of an
ORDER BY id DESC LIMIT 1 backward scan can only match at the same row or an earlier one, so the query never
examines more rows than before.
Raw SQL
thread_ts is the checkpoint ancestor chain, so the IN list holds one entry per checkpoint in the session.
Both queries below use a 300-entry list, truncated here for readability.
Before:
SELECT
"p_duo_workflows_checkpoint_blobs".*
FROM
"p_duo_workflows_checkpoint_blobs"
WHERE
"p_duo_workflows_checkpoint_blobs"."workflow_id" = 42
AND "p_duo_workflows_checkpoint_blobs"."workflow_created_at" = '2026-09-20 12:00:00+00'
AND "p_duo_workflows_checkpoint_blobs"."project_id" = 278964
AND "p_duo_workflows_checkpoint_blobs"."thread_ts" IN ('1755770000.000000', '1755770001.000000', ... )
AND "p_duo_workflows_checkpoint_blobs"."channel" = 'ui_chat_log'
AND "p_duo_workflows_checkpoint_blobs"."step_action" = 'conversation'
ORDER BY
"p_duo_workflows_checkpoint_blobs"."id" DESC
LIMIT 1After:
SELECT
"p_duo_workflows_checkpoint_blobs".*
FROM
"p_duo_workflows_checkpoint_blobs"
WHERE
"p_duo_workflows_checkpoint_blobs"."workflow_id" = 42
AND "p_duo_workflows_checkpoint_blobs"."workflow_created_at" = '2026-09-20 12:00:00+00'
AND "p_duo_workflows_checkpoint_blobs"."project_id" = 278964
AND "p_duo_workflows_checkpoint_blobs"."thread_ts" IN ('1755770000.000000', '1755770001.000000', ... )
AND "p_duo_workflows_checkpoint_blobs"."channel" = 'ui_chat_log'
ORDER BY
"p_duo_workflows_checkpoint_blobs"."id" DESC
LIMIT 1Query plans
Correction to an earlier version of this description: it claimed the table holds no production data. That was wrong. The read flags are off, but the write path has been on, and Database Lab shows roughly 28M rows. A Database Lab plan is being sourced; the plans below come from a seeded GDK and stand in until it lands.
Seed: 10 workflows, 300 checkpoints each, 6 channels, 18,000 rows in one daily partition, ANALYZE run before
measuring. The session's newest ui_chat_log blob is a compaction snapshot, which is the case this MR is about,
and the only case where the two queries differ. Production partitions hold far more per day, so treat the
absolute numbers as a lower bound and the direction as the point.
Before:
Limit (cost=1.04..59.42 rows=1 width=369) (actual time=1.142..1.142 rows=1 loops=1)
Buffers: shared hit=938
-> Index Scan Backward using p_duo_workflows_checkpoint_blobs_20260920_pkey on p_duo_workflows_checkpoint_blobs_20260920
(cost=1.04..1694.04 rows=29 width=369) (actual time=1.142..1.142 rows=1 loops=1)
Index Cond: (workflow_created_at = '2026-09-20 12:00:00+00'::timestamp with time zone)
Filter: ((workflow_id = 1057) AND (project_id = 1) AND (channel = 'ui_chat_log'::text)
AND (step_action = 'conversation'::text) AND (thread_ts = ANY (...)))
Rows Removed by Filter: 16223
Buffers: shared hit=938
Planning Time: 0.546 ms
Execution Time: 1.146 msAfter:
Limit (cost=1.04..55.97 rows=1 width=369) (actual time=1.087..1.088 rows=1 loops=1)
Buffers: shared hit=937
-> Index Scan Backward using p_duo_workflows_checkpoint_blobs_20260920_pkey on p_duo_workflows_checkpoint_blobs_20260920
(cost=1.04..1649.04 rows=30 width=369) (actual time=1.087..1.087 rows=1 loops=1)
Index Cond: (workflow_created_at = '2026-09-20 12:00:00+00'::timestamp with time zone)
Filter: ((workflow_id = 1057) AND (project_id = 1) AND (channel = 'ui_chat_log'::text)
AND (thread_ts = ANY (...)))
Rows Removed by Filter: 16205
Buffers: shared hit=937
Planning Time: 0.523 ms
Execution Time: 1.092 msThe new query stops 18 rows earlier and reads one buffer less. That's the whole difference.
Worth a reviewer's eye
Both plans pick a backward scan on the partition primary key and apply workflow_id, project_id, channel, and
thread_ts as a filter, discarding about 16,000 rows to return one. idx_duo_wf_checkpoint_blobs_dedup isn't used.
This shape predates the MR and isn't changed by it, but it's the read path behind dw_read_blobs_graphql, so it may
deserve an index before that flag is enabled at scale. Happy to open a follow-up if you agree it's a problem.