Geo: Switch from uploads table to partitioned upload tables for replication

Everyone can contribute. Help move this issue forward while earning points, leveling up and collecting rewards.

Purpose

This issue is the rollout log for the upload partition switchover. All 23 geo_<name>_upload_replication and geo_<name>_upload_force_primary_checksumming flags name it as their rollout_issue_url, so every chatops flip on staging, staging-ref, and pre posts a note here.

The plan and the individual changes live in the Phase 6a sub-epic &23525 under &20933. This issue does not duplicate them. Removing the legacy replicator, the dual-run plumbing, and the upload_states and file_registry tables is Phase 6b (&23545).

Where this can be tested before release

GitLab.com production has no Geo. Staging has a Geo primary and no secondary, so it can only exercise primary-side checksumming. Staging-ref is the only GitLab-run environment with a primary and a secondary and replication enabled for every replicable type. It is the one place where partition replication and the Event B secondary backfill can be observed before customers get them.

Rollout order

  1. Enable every geo_<name>_upload_force_primary_checksumming flag on staging and staging-ref. Confirm the parity check is green on both.
  2. On staging-ref, enable every geo_<name>_upload_replication flag. This is a no-op for replication while geo_upload_replication is on. All 23 must be on before the next step, because any partition still off when the legacy flag goes off stops replicating and stops backfilling.
  3. On staging-ref, disable geo_upload_replication and geo_upload_force_primary_checksumming. Watch the secondary backfill for at least 24 hours. Confirm registries populate, verification succeeds, and the Geo Sites page is green.
  4. Only then ship Event A in 19.5 and Event B in 19.6.

Re-enabling geo_upload_replication is the kill switch. It restores the legacy path immediately and the legacy consistency worker backfills anything created while it was off.

Risk: the release is the switch for customers

Self-managed and Dedicated instances get this change by upgrading. Event A in 19.5 turns on 23 flags and Event B in 19.6 turns off the legacy one. There is no percentage ramp and no opt-in. Staging-ref is the only secondary that runs it first.

Self-managed administrators can revert on the Rails console by re-enabling geo_upload_replication. Dedicated tenants have no console access, so for them the upgrade is one way unless a GitLab operator intervenes. Any defect that staging-ref does not surface reaches every Dedicated Geo tenant on their 19.6 upgrade. This is the reason the switchover is split across a required stop and why the staging-ref soak above is mandatory, not optional.

Done when

geo_upload_replication is off on staging-ref, all 23 partition replicators are replicating and verifying there, and Event A and Event B have shipped. Close this issue then. Flag removal happens in Phase 6b.

Background

The original switchover plan and its refinement are in the comments below, in particular the 2026-07-06 review of the code paths involved. Discussion of the dashboard prerequisite that became #628122 (closed) is also here.

Edited by Chloe Fons