Resume offline snapshots shard by shard on every sync
What does this MR do and why?
The issue
An offline snapshot sync that runs out of time can't pick up where it left off. It starts over from the first shard on the next run, so a large registry never finishes.
The sync decides how to resume from one run-level flag, full_sync = checkpoint.first_sync?. On a first sync it resumes per shard; otherwise per archive.
But every shard of a snapshot shares one sequence, so outside a first sync the whole
snapshot becomes a single unit — and a unit is only committed once it's fully ingested.
The budget is 20 minutes for the whole run, shared by every registry in it.
Only the offline path hits this. OfflineV3#data_after re-reads the vendor
full_dataset directory whenever an admin copies in a newer snapshot, which is the
documented way to update an air-gapped instance, so every update after the first one
lands here. Online is fine: PDS only serves /all on a first sync.
The fix
first_sync? was standing in for "this read is a snapshot". That assumption no longer
holds: offline re-reads a snapshot on any run, so a read that isn't the first can't be
taken to be a delta.
The fix moves the flag off the run and onto the data file, where it belongs — a snapshot resumes shard by shard, a delta archive as a whole, and one run can read both.
BaseDataFiletakes asnapshot:keyword, defaultfalse.OfflineV3andPdsmark full-dataset shards withsnapshot: true.new_unit?compares[sequence, chunk]for a shard,[sequence]for a delta file.advance_checkpointsetsfull_sync_target_sequencefrom the file's own flag, and clears it once a later archive commits on top.
Resume itself already worked — the marker keeps first_sync? true, so the next run
goes back to the snapshot and skips the shards it already has. It just never got set.
Note
Nothing changes for PDS. Marking the /all shards is what keeps its existing
per-shard resume working, now that the flag is read off the file. PDS never serves a
snapshot to an already-synced registry, so there's no new behaviour there.
Changes
connector/base_data_file.rb— added asnapshot:keyword (defaultfalse) andsnapshot?.connector/pds.rb—files_fromtakessnapshot:; the/allshards passtrue.connector/offline_v3.rb— full-dataset shards passsnapshot: true.v3_sync_service.rb— addedunit_key; dropped thefull_syncargument fromnew_unit?,advance_checkpointandfinalize_stream; the loop trackspending_fileinstead ofpending_sequence/pending_chunk.v3_sync_service_shared_examples.rb— added asnapshot_filehelper and three contexts; delta assertions now expectfull_sync_target_sequence: nil.
References
How to set up and validate locally
-
Run the specs:
bundle exec rspec ee/spec/services/package_metadata/ ee/spec/lib/gitlab/package_metadata/
New coverage sits in the shared examples both v3 datasets use: a snapshot arriving after the first sync and interrupted mid-way, a snapshot followed by a delta in one stream, and that same stream interrupted at the delta.