Shard large /delta responses (same large-file risk as /all)
Problem
/all is sharded because a single snapshot is too large for one Sidekiq job to fetch, decompress, and ingest within the job-duration limit, and a .tar.zst decompresses to ~13× its compressed size. The connector handles this in full_dataset_files: PDS returns shards (one signed URL per shard), and the sync ingests them shard-by-shard, resuming from the checkpoint chunk if interrupted.
/delta has no equivalent. Both MalwarePds#delta_files (per-registry) and MalwarePds#delta_files_for (bulk multi-registry) treat each delta as a single flat .tar.zst NDJSON archive — there is no shard dimension. That is fine for the steady-state case (a few minutes of new advisories per 5-minute cycle), but a delta can grow to the same size class as an /all shard when:
- an instance has been offline long enough that its cursor is far behind (a long catch-up delta), or
- a single window produces a large burst of advisories (for example a large malware campaign day — see the keyv/cacheable namespace compromise, where a single day added dozens of npm advisories).
When that happens, a large delta archive hits the same OOM / job-duration risks that motivated sharding /all, except there is currently no resume path: an interrupted large delta is re-fetched from scratch on the next run rather than resumed.
Proposal
Give /delta the same sharded shape as /all, end to end:
- PDS side: shard large
/deltaresponses — return per-shard signed URLs for a delta (mirroring the/allshardsshape), so no single archive exceeds the size/duration a single job can safely handle. Applies to both the single-registry and the bulk multi-registry (purl_types-keyed) responses. - Connector: consume sharded deltas the way
full_dataset_filesconsumes/allshards — iterate shards in order, and resume from the checkpointchunkwithin a deltasequenceso an interrupted large delta continues rather than restarting. - Preserve the existing per-archive checkpoint semantics: advance the checkpoint only once a full delta (all its shards) is ingested, so an interrupt mid-delta leaves the cursor at the last fully-ingested delta.
- Specs for: sharded-delta parsing, checkpoint resume by chunk within a delta, and the bulk multi-registry path returning sharded entries per registry.
Acceptance criteria
- A large
/delta(catch-up or burst) is fetched and ingested as shards, none of which individually risks the job-duration limit or OOM. - An interrupted large delta resumes from the last ingested shard on the next run instead of re-fetching the whole delta.
Related
- Builds on the bulk multi-registry
/deltafrom #607448 (closed). - Complements #602885 — streaming decompression bounds memory within one archive; sharding bounds the size and duration per fetch. Both are needed for large-file safety.
- Requires a corresponding PDS-side change to the
/deltaresponse shape (PDS tracking to be linked).