Reject offline malware archives whose purl_type mismatches the directory

What does this MR do?

Fixes an offline (air-gapped) malware advisory sync bug where a mislabelled or misplaced vendor archive advances the wrong registry's checkpoint. Offline keys the checkpoint by the directory it reads (v3/<purl_type>/full_dataset/) but takes each row's purl_type from the payload; nothing reconciled them, so an npm directory full of nuget records advanced the npm cursor while writing nuget data — and npm then resolved first_sync? == false and never fetched its snapshot again, with the run logging success.

Why

In an air-gapped environment there is no upstream to compare against, so a registry that is permanently stuck while logging success is the hardest kind of failure to diagnose. This is the same class as #607059 (closed) (checkpoint advanced without the data being persisted). Offline malware sync is not yet available to anyone, so this can be fixed before it ships.

How

In MalwareAdvisorySyncService#ingest_file, for offline syncs, split off records whose purl_type does not match the sync's registry:

  • Skip the mismatched records — they are not ingested under the wrong registry.
  • Log the mismatch loudly (phase: purl_type_mismatch, with expected_purl_type, found_purl_types, skipped, and sample advisory_xids).
  • Do not advance the checkpoint for a slice that contained a mismatch (persisted_all = false), so the mislabelled archive is re-read next run rather than silently skipped — enforcing the #607059 (closed) invariant (do not advance a checkpoint that persisted no rows).

PDS is unaffected: it requests one registry and the response only contains that registry (the check is gated on storage_type == :offline).

Testing

Unit spec added to ee/spec/services/package_metadata/malware_advisory_sync_service_spec.rb: an offline archive whose payload purl_type differs from its directory is skipped, logged, and does not advance the checkpoint.

Local testing

Offline directory/payload mismatch — before/after

Reproduced on a local GDK whose offline v3/npm/full_dataset/ is mislabelled — its records are all purl_type: nuget. Reset the npm checkpoint and run a full offline sync:

# gitlab-rails runner
ENV['PM_SYNC_IN_DEV'] = 'true'
config = PackageMetadata::SyncConfiguration.configs_for('malware_advisories').find { |c| c.purl_type.to_s == 'npm' }
cp = PackageMetadata::Checkpoint.for_malware_advisories.find_by!(purl_type: :npm)
cp.update!(sequence: 0, chunk: 0, full_sync_target_sequence: nil)
never = Object.new; never.define_singleton_method(:stop?) { false }
PackageMetadata::MalwareAdvisorySyncService.new(config, never, checkpoint: cp).execute
cp.reload
puts "npm checkpoint: seq=#{cp.sequence} chunk=#{cp.chunk} first_sync?=#{cp.first_sync?}"

Before (master) — the npm cursor is earned from nuget data, silently:

npm checkpoint: seq=1783005651 chunk=63 first_sync?=false
{"phase":"completed","purl_type":"npm","message":"Malware advisory full sync completed"}

After (this MR) — the checkpoint is not advanced, and the mismatch is logged per shard:

npm checkpoint: seq=0 chunk=0 first_sync?=true
{"phase":"purl_type_mismatch","expected_purl_type":"npm","found_purl_types":["nuget"],"skipped":12,
 "message":"Malware advisory sync: archive purl_type does not match its directory"}

npm stays first_sync? == true and re-reads /all next run instead of recording a completed sync it did not do.

References

Merge request reports

Loading
Loading