Geo: bulk_import_export_upload_upload_states fails NOT NULL on project_id for group-level Direct Transfer exports

Summary

The Geo verification state table bulk_import_export_upload_upload_states cannot store state for group-level Direct Transfer export uploads. Every attempt raises:

PG::NotNullViolation: ERROR:  null value in column "project_id" of relation
"bulk_import_export_upload_upload_states" violates not-null constraint (ActiveRecord::NotNullViolation)
DETAIL:  Failing row contains (130272, null, null, null, 35180, null, 0, 0, null, null).

Seen on staging-ref via Sidekiq/Geo::VerificationStateBackfillWorker (386 events and climbing): Sentry issue

Cause

BulkImports::Export (and therefore its ExportUpload files) is owned by either a project or a group — exactly one:

# app/models/bulk_imports/export.rb
validates :project, presence: true, unless: :group
validates :group, presence: true, unless: :project

The parent partition table bulk_import_export_upload_uploads models this correctly: nullable project_id/namespace_id with CHECK (num_nonnulls(namespace_id, project_id) = 1).

But the state table created in !229718 (merged) assumes project-only ownership:

  • project_id bigint NOT NULL, no namespace_id column (migration)
  • a single sharding-key trigger copying project_id from the parent row (migration)

When a group-level export exists, the parent row's project_id is NULL, the trigger copies NULL into the NOT NULL column, and the insert fails. This breaks both:

  1. the after_save verification-details callback on every save of such an upload, and
  2. Geo::VerificationStateBackfillWorker — worse here, because the batch insert_all is all-or-nothing: one group-owned upload wedges state creation for every record in its batch range. The batcher never advances and the cron retries (and fails) every minute — hence the Sentry event volume.

Introduced in !229718 (merged) (19.0). Impact is currently limited because the geo_bulk_import_export_upload_upload_replication / geo_bulk_import_export_upload_upload_force_primary_checksumming ops flags default to off; staging-ref has them enabled, and this will break verification anywhere they're enabled.

Reference implementations that got it right

Two sibling state tables from &20933 handle multi-owner replicables correctly and should be used as the template:

  • !229562 (merged) — import_export_upload_upload_states: nullable project_id + namespace_id, FKs to both, add_multi_column_not_null_constraint, and two sharding-key triggers (migration)
  • dependency_list_export_upload_states: same pattern with three owner columns

An audit of all 23 upload state tables from &20933 found this is the only broken one (all others have a guaranteed single owner). One adjacent observation: personal_snippet_upload_states.organization_id NOT NULL relies on app-level-only validation of snippets.organization_id — not a bug today, but worth keeping in mind.

Proposed fix (single MR, no data cleanup needed)

The NOT NULL constraint rejected every bad insert, so no invalid rows exist and no backfill is required. The wedged backfill batch self-heals on its next cron run after the migration.

  1. add_column :bulk_import_export_upload_upload_states, :namespace_id, :bigint + index + FK to namespaces (on_delete: :cascade)
  2. change_column_null :bulk_import_export_upload_upload_states, :project_id, true (Note: staging-ref / staging may now have data in this table due to FF enablement so we may need to check and manually remove this data first on these environments)
  3. add_multi_column_not_null_constraint(:bulk_import_export_upload_upload_states, :project_id, :namespace_id)
  4. install_sharding_key_assignment_trigger for namespace_id (mirroring the existing project_id trigger)
  5. Update db/docs/bulk_import_export_upload_upload_states.yml
  6. Geo::BulkImportExportUploadUpload: add namespace_id_in scope and include namespace-owned records in selective_sync_scope (mirror Geo::ImportExportUploadUpload) — group-owned uploads are currently also silently excluded from selective sync

Why tests never caught this

The geo_bulk_import_export_upload_upload factory hardcodes a project-owned export (project { create(:project) } transient), so no Geo spec ever creates a group-owned upload — every test insert had project_id populated by the trigger. The CI sharding-key spec only verifies that a sharding key column exists and references projects/namespaces/organizations, which project_id NOT NULL satisfies. The reference MR !229562 (merged) has the same test gap; it passes because its schema happens to be correct.

Recommendation: add coverage

As part of the fix (written first, to demonstrate red → green):

  1. Model spec: a group-owned upload can persist verification details (reproduces the callback failure)
  2. Backfill spec: create_verification_details_for with a group-owned record (reproduces the exact worker failure)
  3. Selective-sync spec: group-owned uploads are included in selective_sync_scope for namespace-based selective sync (currently silently excluded — wrong results, not an error)
  4. Factory: add a :group_owned trait/transient to geo_bulk_import_export_upload_upload (and consider the same for the import_export factory, which shares the gap)
Edited by 🤖 GitLab Bot 🤖