Resume offline snapshots shard by shard on every sync

What does this MR do and why?

The issue

An offline snapshot sync that runs out of time can't pick up where it left off. It starts over from the first shard on the next run, so a large registry never finishes.

The sync decides how to resume from one run-level flag, full_sync = checkpoint.first_sync?. On a first sync it resumes per shard; otherwise per archive. But every shard of a snapshot shares one sequence, so outside a first sync the whole snapshot becomes a single unit — and a unit is only committed once it's fully ingested. The budget is 20 minutes for the whole run, shared by every registry in it.

Only the offline path hits this. OfflineV3#data_after re-reads the vendor full_dataset directory whenever an admin copies in a newer snapshot, which is the documented way to update an air-gapped instance, so every update after the first one lands here. Online is fine: PDS only serves /all on a first sync.

The fix

first_sync? was standing in for "this read is a snapshot". That assumption no longer holds: offline re-reads a snapshot on any run, so a read that isn't the first can't be taken to be a delta.

The fix moves the flag off the run and onto the data file, where it belongs — a snapshot resumes shard by shard, a delta archive as a whole, and one run can read both.

  • BaseDataFile takes a snapshot: keyword, default false.
  • OfflineV3 and Pds mark full-dataset shards with snapshot: true.
  • new_unit? compares [sequence, chunk] for a shard, [sequence] for a delta file.
  • advance_checkpoint sets full_sync_target_sequence from the file's own flag, and clears it once a later archive commits on top.

Resume itself already worked — the marker keeps first_sync? true, so the next run goes back to the snapshot and skips the shards it already has. It just never got set.

Note

Nothing changes for PDS. Marking the /all shards is what keeps its existing per-shard resume working, now that the flag is read off the file. PDS never serves a snapshot to an already-synced registry, so there's no new behaviour there.

Changes

  • connector/base_data_file.rb — added a snapshot: keyword (default false) and snapshot?.
  • connector/pds.rb — files_from takes snapshot:; the /all shards pass true.
  • connector/offline_v3.rb — full-dataset shards pass snapshot: true.
  • v3_sync_service.rb — added unit_key; dropped the full_sync argument from new_unit?, advance_checkpoint and finalize_stream; the loop tracks pending_file instead of pending_sequence/pending_chunk.
  • v3_sync_service_shared_examples.rb — added a snapshot_file helper and three contexts; delta assertions now expect full_sync_target_sequence: nil.

References

How to set up and validate locally

  1. Run the specs:

    bundle exec rspec ee/spec/services/package_metadata/ ee/spec/lib/gitlab/package_metadata/

New coverage sits in the shared examples both v3 datasets use: a snapshot arriving after the first sync and interrupted mid-way, a snapshot followed by a delta in one stream, and that same stream interrupted at the delta.

Edited by Orin Naaman

Merge request reports

Loading
Loading