Read legacy checkpoints when a header records no channel membership
Problem
Duo Agent Platform flows store checkpoints two ways: a legacy full row per checkpoint, and a new header row plus append-only blobs. Headers have been written since 2026-07-30, but the header's channel_keys column only exists since !250189 (merged) deployed, so every header written before that has NULL. Blobs cannot express a channel deletion, so a NULL membership cannot drive the fold filter that #613975 (closed) adds, and it cannot be backfilled.
What changed
Workflow#legacy_checkpoint_fallback?treats a NULLchannel_keyson the newest header as "not incremental", falling back to the legacy table.Workflow#read_from_blobs?is the new shared read gate. GraphQL, the internal checkpoint list, notifications and the trace download all read through it.- The internal
by_thread_tsendpoint checks its own header and serves the legacy row (or 404s) when the membership is NULL. Checkpoint.latest_for_headerfinds the matching legacy row bythread_ts, bounded by the existing clock-skew window.
Why this is safe
The write flag duo_workflow_write_incremental_only has never been enabled, so every pre-column header still has its full legacy row behind it, and channel_keys = [] (a valid empty value) keeps the blob path. During the fallback window a session's GraphQL message history is truncated at the last compaction, which matches current behavior for a workflow with incremental checkpoints disabled.
Pre-column headers age out with their 30-day partition, and #611971 removes the fallback once they do.
In incremental-only mode the gateway drops the field that CreateCheckpointService derives channel_keys from, so those headers would get NULL with no legacy row behind them. #617794 (closed) and gitlab-org/modelops/applied-ml/code-suggestions/ai-assist#2745 (closed) must both deploy before duo_workflow_write_incremental_only is enabled.
Why the read cannot flip mid-request
The read decision happens once per request, before any row loads: the WorkflowPresenter GraphQL methods, the internal checkpoint list, #latest_readable_checkpoint, and the trace read each pick the table once and reuse it. The internal by_thread_ts endpoint checks only the one header it serves.
One caller asks per event: WorkflowCheckpointEventPresenter calls event.workflow.reconstruct_from_blobs_for_graphql? for each node in a page. That could split a page across two decisions, but #legacy_checkpoint_fallback? memoizes per instance, and event.workflow resolves to the same Workflow instance through inverse_of. A spec in workflow_presenter_spec.rb asserts that no extra header query runs when a page is presented, so it fails if that shared-instance behavior breaks.
The residual risk is narrow: a NULL-membership header would have to land between two queries on two different Workflow instances in the same request. That needs a writer still emitting NULL, which means a pre-column row, a malformed payload, or incremental-only mode.
Query plans for database review
Plan 1, the gate's newest-header read (#latest_checkpoint_header):
SELECT "p_duo_workflows_checkpoint_headers".*
FROM "p_duo_workflows_checkpoint_headers"
WHERE "workflow_created_at" = '2026-07-27 11:58:06.045052+00'
AND "workflow_id" = 624
AND "checkpoint_ns" IS NULL
ORDER BY "thread_ts" DESC, "id" DESC
LIMIT 1; Limit (cost=0.14..18.64 rows=1 width=1172) (actual time=0.016..0.016 rows=1 loops=1)
Buffers: shared hit=5
-> Index Scan Backward using index_3052e17b3e on p_duo_workflows_checkpoint_headers_20260727 p_duo_workflows_checkpoint_headers (cost=0.14..18.64 rows=1 width=1172) (actual time=0.015..0.015 rows=1 loops=1)
Index Cond: (workflow_id = 624)
Filter: ((checkpoint_ns IS NULL) AND (workflow_created_at = '2026-07-27 11:58:06.045052+00'::timestamp with time zone))
Buffers: shared hit=5
Planning Time: 5.220 ms
Execution Time: 0.028 msPlan 2, the by_thread_ts fallback lookup (Checkpoint.latest_for_header):
SELECT "p_duo_workflows_checkpoints".*
FROM "p_duo_workflows_checkpoints"
WHERE "workflow_id" = 1027
AND "created_at" BETWEEN '2026-08-12 12:55:25.961397+00' AND '2026-08-12 14:55:25.961397+00'
AND "thread_ts" = '1f196557-82a6-617c-8007-5b330aff40b0'
ORDER BY "id" DESC
LIMIT 1; Limit (cost=2.18..2.19 rows=1 width=244) (actual time=0.031..0.031 rows=1 loops=1)
Buffers: shared hit=8
-> Sort (cost=2.18..2.19 rows=1 width=244) (actual time=0.030..0.031 rows=1 loops=1)
Sort Key: p_duo_workflows_checkpoints.id DESC
Sort Method: quicksort Memory: 25kB
-> Index Scan using p_duo_workflows_checkpoints_20260812_workflow_id_thread_ts_idx on p_duo_workflows_checkpoints_20260812 p_duo_workflows_checkpoints (cost=0.15..2.17 rows=1 width=244) (actual time=0.022..0.022 rows=1 loops=1)
Index Cond: ((workflow_id = 1027) AND (thread_ts = '1f196557-82a6-617c-8007-5b330aff40b0'::text))
Filter: ((created_at >= '2026-08-12 12:55:25.961397+00'::timestamp with time zone) AND (created_at <= '2026-08-12 14:55:25.961397+00'::timestamp with time zone))
Planning Time: 0.799 ms
Execution Time: 0.044 msBoth plans come from a local GDK database with 58 daily partitions on p_duo_workflows_checkpoints. Plan 1 uses enable_seqscan = off, because the local partition holds 8 rows.
Test coverage
workflow_spec.rbcovers#read_from_blobs?across NULL, empty and populatedchannel_keys, an older header without a membership, no headers, and the read flag off.checkpoint_spec.rbcovers.latest_for_header: the matching row, the latest row after a re-send, an out-of-window row ignored, and nil when nothing matches.workflows_internal_spec.rbcoversby_thread_tsserving the legacy row for a NULL-membership header, and 404 when that row is gone.
Follow-ups
- Add
.latest_for_headerto the restricted-methods list of the legacy-checkpoint-read cop from !249745 (merged) once it lands. - Wrap the
by_thread_tsfallback query in thewith_legacy_readsuppression once !250618 (closed) lands.