Persist channel membership on the checkpoint header (lost under duo_workflow_write_incremental_only)
Problem
Blob reconstruction cannot express channel deletion, so it cannot reproduce the exact set of channels a checkpoint had. This surfaced as QA scenario 2 on !247134 (merged) (note): the reconstructed channel_values drops LangGraph's branch:to:* channels (the scheduler's run queue), so a resumed graph sees no armed tasks and silently terminates.
Root cause is a wrong assumption: channel membership (which channels exist at a given checkpoint) was implicitly recoverable from the full checkpoint row / compressed_checkpoint. Under duo_workflow_write_incremental_only that full row is no longer written, so membership is lost:
- Blobs are an append-only per-group log. The reconstructed key set is the union of every channel ever blobbed in the group — there is no way to say "this channel is gone." Channels are deleted on ~85% of steps (
apply_writesclears armed ephemeral channels each step), so stale keys accumulate. - Scalar channels (
branch:to:*,approval,goal, …) are not blobbed at all today, so their presence is lost entirely.
channel_versions (already stored in the header checkpoint JSONB) cannot stand in for membership: it retains versions for consumed channels (e.g. 10 channel_versions keys vs 6 live channel_values keys at the boundary checkpoint).
What (write path / this issue)
Persist the live channel membership — channel_values.keys — on the checkpoint header at write time, so the read path can select exactly the channels that existed.
The keys are already in hand at write time. CreateCheckpointService#write_checkpoint_header (ee/app/services/ai/duo_workflows/create_checkpoint_service.rb:132) currently strips channel_values for the slim header:
checkpoint: checkpoint.checkpoint.except('channel_values', :channel_values),Capture checkpoint.checkpoint['channel_values']&.keys there and store it.
Table change
Add a column to p_duo_workflows_checkpoint_headers for the membership list (a handful of short strings), e.g. channel_keys text[] (or a JSONB array). The table is range-partitioned by workflow_created_at, so use the partitioned-table migration helpers. Include up/down migration output and the db/structure.sql regen (via scripts/regenerate-schema).
Populate it in write_checkpoint_header and cover it in create_checkpoint_service_spec + the request spec.
Depends on / related (not this issue)
- Read/reconstruction (!247134 (merged) / read epic gitlab-org/gitlab&22939): fold the blobs, then select only the keys the header declares. Deletions then cost nothing and
__start__stops resurrecting. - Gateway (ai-assist !6363 (merged) and follow-up): blob every channel regardless of type (drop the scalar allowlist) so
branch:to:*content (value=null) is present; otherwise the header lists a key with no blob behind it.
These three land together to fix the bug; this issue is the header/table part.
Traps
- Do not reuse
channel_versionsas the membership list (see above). - If the team later prefers tombstone blobs (
step_action: "delete") over a header membership list, the delete sentinel must not benil—branch:to:X = nullis a legitimate live value.
Related
- MR note with full analysis: !247134 (comment 3652137773)
- Write-only rollout FF: #607029
- Sibling: #613454 (closed) (stop sending
compressed_checkpointunder write-only)