Advisory-driven Continuous Vulnerability Scanning generates sustained WAL pressure on the sec database

Everyone can contribute. Help move this issue forward while earning points, leveling up and collecting rewards.

We are seeing recurring PrimaryDatabaseWALGenerationSaturationSpike alerts on the patroni-sec cluster (gitlabhq_production_sec), the second occurrence in under a month. After digging through production logs, the spikes trace back to advisory-driven Continuous Vulnerability Scanning, and the load profile looks likely to grow as we expand tracking with Vulnerabilities Across Contexts.

This was actually anticipated in #437162 (closed), which raised concern about "the bulk inserts done by the IngestCvsSliceService... If we ingest too many advisories at once (worker creation is currently unbounded), and they happen to match multiple projects... resource contention, and an eventual degradation." We now have production evidence that this is happening, so I am opening a dedicated issue to track a fix.

Note: supporting screenshots (Grafana alert and Kibana log views) are kept in an internal comment on this issue, as they contain internal hostnames and IP addresses. GitLab team members can view them there.

What is happening

When an advisory is ingested, PackageMetadata::GlobalAdvisoryScanWorker runs one job per advisory and scans every project whose SBOM contains the affected package. The fan-out lives inside the worker: Gitlab::VulnerabilityScanning::AdvisoryScanner#scan_projects_for walks all matching Sbom::Occurrence rows across all projects, then bulk upserts findings through Security::Ingestion::IngestCvsSliceService into vulnerability_occurrences, vulnerabilities, and vulnerability_reads on the sec primary. Those bulk writes are what generate the WAL.

Production evidence

These are from the gprd logs over a 3 hour window covering one of the spikes (2026-06-12).

On the sec primary, the slow write statements (INSERT/UPDATE/DELETE on patroni-sec*) broken down by the originating worker are bursty and dominated by security ingestion (Kibana: sec primary writes by worker (3h)). The two dominant writers during the spike were Security::StoreSecurityReportsByProjectWorker (~72%) and PackageMetadata::GlobalAdvisoryScanWorker (~22%), both upserting the vulnerability tables.

Looking at the advisory scan jobs themselves, there were 394 completed GlobalAdvisoryScanWorker runs in that 3 hour window, cron-triggered from the advisory feed (Kibana: GlobalAdvisoryScanWorker completions (3h)). A single representative job ran for 517 seconds and issued 17,367 queries against the sec database (83 seconds of sec DB time), processing a single advisory_id. All the advisory-scan writes captured on the sec primary in the window shared one correlation id, so they are one continuous fan-out rather than many small independent ones. At hundreds of these jobs every few hours, each touching tens of thousands of sec rows, the cumulative upsert volume keeps WAL generation elevated and occasionally pushes it into saturation territory, where replicas start to lag and read traffic gets pinned to the primary.

Why this likely gets worse with VAC

Today the scanner fans out across Sbom::Occurrence rows on the default branch. VAC expands tracking to tags and non-default branches, which multiplies the number of occurrences and findings per project. Since the advisory scanner re-evaluates all matching occurrences, the per-advisory write volume scales with the number of tracked contexts. As VAC rolls out beyond Customer 0, the same advisory feed will translate into proportionally more sec writes, so the WAL pressure we are seeing now is probably a floor, not a ceiling.

Possible directions

My strongest suggestion is to reuse the database health backpressure we already have. Gitlab::Database::HealthStatus (health_status.rb) evaluates indicators including a WAL one (WriteAheadLog) and emits a stop signal under pressure. It already throttles Sidekiq jobs: any worker that declares defer_on_database_health_signal? gets deferred when the database signals stop, via skip_jobs.rb. Batched background migrations already lean on this same framework to pause themselves under WAL pressure (see Throttling batched migrations). Having GlobalAdvisoryScanWorker (and/or the CVS ingestion path) opt into the same signal would let advisory scanning yield to database health automatically rather than competing with it.

Looking further ahead, the newer Batched Background Operations framework (background_operations.md) supports recurring, cron-scheduled work and consumes the very same HealthStatus checks for backpressure, so it could be a natural longer-term home for advisory scanning. It is still flagged as experimental (the docs ask you to reach out to #g_database_architecture before adopting), so I mention it as a direction to consider rather than something to pick up today.

Other options worth weighing, not mutually exclusive:

  • Smooth out job creation. The worker has concurrency_limit 10 and urgency :throttled, but advisory ingestion can still enqueue a large batch at once. Spacing out enqueues or capping advisories processed per interval would flatten the write bursts.
  • Smaller or rate-limited ingestion slices. Reducing the IngestCvsSliceService batch size or pacing the slices would lower the instantaneous WAL rate per job.
  • Off-peak scheduling. If freshness allows, shifting bulk advisory re-scans to lower-traffic windows would reduce contention with live read traffic.
  • Anything else the team sees as more idiomatic for this path.

References

cc @hacks4oats @nilieskou @cwidstrom @ryaanwells

Edited by 🤖 GitLab Bot 🤖