Batch firstCheckpoint across GraphQL pages

What does this MR do and why?

The IDE session list query (duoWorkflowWorkflows) selects firstCheckpoint on every row of a page. On the incremental-blob read path (feature flag dw_read_blobs_graphql), each row ran two header queries: the read gate's newest-header lookup and the oldest-header lookup that serves firstCheckpoint. A 50-row page ran about 4 checkpoint queries per row. See the gprd evidence in the linked work item.

This MR batches firstCheckpoint at the page level. It follows the pattern that !253884 (merged) introduced for latestCheckpoint and duoMessages.

Changes:

  1. New scope Ai::DuoWorkflows::CheckpointHeader.earliest_per_workflow. It uses DISTINCT ON (workflow_id) ordered by workflow_id, thread_ts ASC, id ASC. It mirrors latest_per_workflow in ascending order. It carries no partition bound. Callers compose it with a workflow_created_at filter.
  2. New method Ai::DuoWorkflows::Workflow.prime_earliest_checkpoint_headers(workflows, checkpoint_ns:). It loads the oldest header per workflow in one partition-pruned query and stores it on each workflow instance. Workflow#earliest_checkpoint_header is now memoized per lineage (checkpoint_ns), so a primed value is reused. Only workflows for which incremental_blob_gate.graphql_candidate? is true join the batch. Legacy-path workflows keep their existing per-row checkpoints.earliest query.
  3. Types::Ai::DuoWorkflows::WorkflowType#first_checkpoint is now a BatchLoader::GraphQL loader keyed by [presenter, checkpoint_ns]. It first primes the newest headers for the page (the read gate IncrementalBlobGate#for_graphql? reads them through Workflow#legacy_checkpoint_fallback?). It then primes the oldest header per lineage. It then resolves each row through the existing presenter method WorkflowPresenter#first_checkpoint.

Result: on a page that selects only firstCheckpoint, checkpoint-header queries drop from 2 per row to 2 per page. The new request spec asserts exactly 2 queries on p_duo_workflows_checkpoint_headers and 0 on p_duo_workflows_checkpoint_blobs for a 3-row page. With the loader disabled, the same spec sees 6 header queries.

Not in scope: the deprecated DuoWorkflowEvent.checkpoint field (deprecated since 18.7). When a client selects firstCheckpoint { checkpoint }, Rails still reconstructs the checkpoint's channel_values from blobs per row. Clients should stop selecting this field. A client issue is filed at gitlab-org/editor-extensions/gitlab-lsp#3009 (closed).

This MR makes no GraphQL schema change, no feature flag change, no migration, and no change to field values.

References

Resolves the backend part (item 1) of #628051 (closed).

Sibling of #627867 (closed).

Database

This MR adds one new read-only query: the earliest_per_workflow scope composed with workflow_id IN (...), workflow_created_at IN (...), and checkpoint_ns IS NULL. The query prunes to one daily partition by workflow_created_at. The existing per-partition index on (workflow_id, thread_ts, id) serves it.

Query plan

Plan from postgres.ai Database Lab (gitlab-production-main), for a real 20-workflow page (120 header rows), cold cache. Total time: 21.8 ms (planning 2.9 ms, execution 18.8 ms, 38 shared hits, 42 reads).

SQL
SELECT DISTINCT ON (workflow_id) "p_duo_workflows_checkpoint_headers".* FROM "p_duo_workflows_checkpoint_headers" WHERE "p_duo_workflows_checkpoint_headers"."workflow_id" IN (6799229,6799230,6799231,6799232,6799233,6799234,6799235,6799236,6799237,6799238,6799239,6799240,6799241,6799242,6799243,6799244,6799245,6799246,6799247,6799248) AND "p_duo_workflows_checkpoint_headers"."workflow_created_at" IN ('2026-09-01 00:00:12.266806+00','2026-09-01 00:00:15.319162+00','2026-09-01 00:00:20.689198+00','2026-09-01 00:00:32.107674+00','2026-09-01 00:00:37.011618+00','2026-09-01 00:00:46.921664+00','2026-09-01 00:00:50.119548+00','2026-09-01 00:00:58.066413+00','2026-09-01 00:01:01.154622+00','2026-09-01 00:01:03.797742+00','2026-09-01 00:01:11.14093+00','2026-09-01 00:01:21.289201+00','2026-09-01 00:01:22.443609+00','2026-09-01 00:01:28.820395+00','2026-09-01 00:01:31.116024+00','2026-09-01 00:01:31.280698+00','2026-09-01 00:01:35.642479+00','2026-09-01 00:01:39.674303+00','2026-09-01 00:01:43.304268+00','2026-09-01 00:01:45.908476+00') AND "p_duo_workflows_checkpoint_headers"."checkpoint_ns" IS NULL ORDER BY "p_duo_workflows_checkpoint_headers"."workflow_id" ASC, "p_duo_workflows_checkpoint_headers"."thread_ts" ASC, "p_duo_workflows_checkpoint_headers"."id" ASC
Plan
Unique  (cost=0.47..317.76 rows=1 width=1311) (actual time=4.025..18.770 rows=20 loops=1)
  Buffers: shared hit=38 read=42
  I/O Timings: read=18.456 write=0.000
  ->  Index Scan using index_b8ef6e7d36 on gitlab_partitions_dynamic.p_duo_workflows_checkpoint_headers_20260901 p_duo_workflows_checkpoint_headers  (cost=0.47..317.76 rows=1 width=1311) (actual time=4.023..18.744 rows=120 loops=1)
        Index Cond: (p_duo_workflows_checkpoint_headers.workflow_id = ANY ('{6799229,6799230,6799231,6799232,6799233,6799234,6799235,6799236,6799237,6799238,6799239,6799240,6799241,6799242,6799243,6799244,6799245,6799246,6799247,6799248}'::bigint[]))
        Filter: ((p_duo_workflows_checkpoint_headers.checkpoint_ns IS NULL) AND (p_duo_workflows_checkpoint_headers.workflow_created_at = ANY ('{"2026-09-01 00:00:12.266806+00","2026-09-01 00:00:15.319162+00","2026-09-01 00:00:20.689198+00","2026-09-01 00:00:32.107674+00","2026-09-01 00:00:37.011618+00","2026-09-01 00:00:46.921664+00","2026-09-01 00:00:50.119548+00","2026-09-01 00:00:58.066413+00","2026-09-01 00:01:01.154622+00","2026-09-01 00:01:03.797742+00","2026-09-01 00:01:11.14093+00","2026-09-01 00:01:21.289201+00","2026-09-01 00:01:22.443609+00","2026-09-01 00:01:28.820395+00","2026-09-01 00:01:31.116024+00","2026-09-01 00:01:31.280698+00","2026-09-01 00:01:35.642479+00","2026-09-01 00:01:39.674303+00","2026-09-01 00:01:43.304268+00","2026-09-01 00:01:45.908476+00"}'::timestamp with time zone[])))
        Buffers: shared hit=38 read=42
        I/O Timings: read=18.456 write=0.000
Settings: seq_page_cost = '4', effective_cache_size = '472585MB', jit = 'off', random_page_cost = '1.5', work_mem = '230MB'

Time: 21.773 ms
  - planning: 2.944 ms
  - execution: 18.829 ms
    - I/O read: 18.456 ms
    - I/O write: 0.000 ms

Shared buffers:
  - hits: 38 (~304.00 KiB) from the buffer pool
  - reads: 42 (~336.00 KiB) from the OS file cache, including disk I/O
  - dirtied: 0
  - writes: 0

index_b8ef6e7d36 is the partition-local name of the (workflow_id, thread_ts, id) index on p_duo_workflows_checkpoint_headers_20260901.

How to set up and validate locally

  1. Enable the duo_workflow_read_incremental_checkpoints and dw_read_blobs_graphql feature flags for a project.
  2. Create several sessions with incremental_checkpoints_enabled: true and checkpoint headers.
  3. Run a duoWorkflowWorkflows query that selects nodes { id firstCheckpoint { threadTs } }.
  4. Check the query count with ActiveRecord::QueryRecorder or the performance bar. Expect two queries on p_duo_workflows_checkpoint_headers, regardless of row count.

Or run:

bundle exec rspec ee/spec/requests/api/graphql/ai/duo_workflows/workflows_spec.rb -e firstCheckpoint

MR acceptance checklist

This checklist encourages us to confirm any changes have been analyzed to reduce risks in quality, performance, reliability, security, and maintainability. See the acceptance checklist for details.

  • I have evaluated the MR acceptance checklist for this MR.
Edited by Eduardo Bonet

Merge request reports

Loading
Loading