Major version upgrade of a citus SGShardedCluster loses the citus metadata
Summary
A major version upgrade of a citus SGShardedCluster destroys the citus metadata. The cluster
comes up on the target PostgreSQL version, but it is not a citus cluster anymore: no nodes, no
distributed tables, no shards.
pg_upgrade migrates the schema and the data of user tables, but not the data of the tables owned
by an extension: it dumps the old cluster with pg_dump --schema-only --binary-upgrade. Every
citus catalog is therefore empty on the upgraded cluster, pg_dist_node and pg_dist_local_group
included.
This is what citus_prepare_pg_upgrade() and citus_finish_pg_upgrade() exist for. The former
copies the citus catalogs into regular public.pg_dist_* tables that pg_upgrade does carry over,
the latter restores them and drops the copies. They must be called on the coordinator and on every
worker, around pg_upgrade. Neither is called anywhere:
stackgres-k8s/src/operator/src/main/resources/templates/major-version-upgrade.sh runs pg_upgrade
without them.
Impact
- The cluster cannot recover on its own, and it cannot be repaired by re-registering the nodes
either.
pg_dist_local_groupis empty too, so citus does not know the coordinator is the coordinator: everycitus_add_nodefails withoperation is not allowed on this node / HINT: Connect to the coordinator and run it again. Patroni retries it for every group every ten seconds, forever. - Distributed tables become plain local tables on the coordinator. The shards are left orphaned on the workers, so the rows they hold are no longer reachable through the coordinator and the rows that were on the coordinator before the table was distributed reappear as the whole table. Any query against them silently returns wrong results instead of failing.
- The reshard, the reference tables and every other citus operation are lost with the metadata.
- It affects every citus
SGShardedClusterthat goes through amajorVersionUpgradeSGShardedDbOps, with and withoutlink.
Proposed resolution
Call SELECT citus_prepare_pg_upgrade() on the primary of the coordinator and of every worker
before the SGDbOps that upgrade them are created, and SELECT citus_finish_pg_upgrade() once all
of them have terminated, from the SGShardedDbOps job (run-sharded-major-version-upgrade.sh).
On rollback the clusters are back on the source postgres version with their citus metadata
untouched, so citus_finish_pg_upgrade() is performed in that case too and only drops the copies
left behind by citus_prepare_pg_upgrade().
The presence of public.pg_dist_node tells whether the metadata has already been saved, so that a
retry of the job does not save the empty catalog of an already upgraded cluster over the copies that
hold the only remaining version of it.
The e2e spec sharded-dbops-major-version-upgrade has to assert that the mock data is still
distributed after the upgrade. It only checks the row count read from the coordinator, which keeps
passing once the distributed table has become a local one, which is why this went unnoticed.