Persist channel membership on the checkpoint header (lost under duo_workflow_write_incremental_only)

Problem

Blob reconstruction cannot express channel deletion, so it cannot reproduce the exact set of channels a checkpoint had. This surfaced as QA scenario 2 on !247134 (merged) (note): the reconstructed channel_values drops LangGraph's branch:to:* channels (the scheduler's run queue), so a resumed graph sees no armed tasks and silently terminates.

Root cause is a wrong assumption: channel membership (which channels exist at a given checkpoint) was implicitly recoverable from the full checkpoint row / compressed_checkpoint. Under duo_workflow_write_incremental_only that full row is no longer written, so membership is lost:

  • Blobs are an append-only per-group log. The reconstructed key set is the union of every channel ever blobbed in the group — there is no way to say "this channel is gone." Channels are deleted on ~85% of steps (apply_writes clears armed ephemeral channels each step), so stale keys accumulate.
  • Scalar channels (branch:to:*, approval, goal, …) are not blobbed at all today, so their presence is lost entirely.

channel_versions (already stored in the header checkpoint JSONB) cannot stand in for membership: it retains versions for consumed channels (e.g. 10 channel_versions keys vs 6 live channel_values keys at the boundary checkpoint).

What (write path / this issue)

Persist the live channel membership — channel_values.keys — on the checkpoint header at write time, so the read path can select exactly the channels that existed.

The keys are already in hand at write time. CreateCheckpointService#write_checkpoint_header (ee/app/services/ai/duo_workflows/create_checkpoint_service.rb:132) currently strips channel_values for the slim header:

checkpoint: checkpoint.checkpoint.except('channel_values', :channel_values),

Capture checkpoint.checkpoint['channel_values']&.keys there and store it.

Table change

Add a column to p_duo_workflows_checkpoint_headers for the membership list (a handful of short strings), e.g. channel_keys text[] (or a JSONB array). The table is range-partitioned by workflow_created_at, so use the partitioned-table migration helpers. Include up/down migration output and the db/structure.sql regen (via scripts/regenerate-schema).

Populate it in write_checkpoint_header and cover it in create_checkpoint_service_spec + the request spec.

  • Read/reconstruction (!247134 (merged) / read epic gitlab-org/gitlab&22939): fold the blobs, then select only the keys the header declares. Deletions then cost nothing and __start__ stops resurrecting.
  • Gateway (ai-assist !6363 (merged) and follow-up): blob every channel regardless of type (drop the scalar allowlist) so branch:to:* content (value=null) is present; otherwise the header lists a key with no blob behind it.

These three land together to fix the bug; this issue is the header/table part.

Traps

  • Do not reuse channel_versions as the membership list (see above).
  • If the team later prefers tombstone blobs (step_action: "delete") over a header membership list, the delete sentinel must not be nil — branch:to:X = null is a legitimate live value.