SGShardedCluster: query routers removed by decreasing queryRouterClusters are never removed from pg_dist_node

Summary

When spec.coordinator.queryRouterClusters of a citus SGShardedCluster is decreased, the query routers that are removed stay registered in pg_dist_node forever. Neither Patroni (that only acts on the Citus groups present in the DCS) nor StackGres ever call citus_remove_node for them.

The stale row is a metadata node (hasmetadata = t) that can not be reached, so Citus marks it metadatasynced = f and, since every metadata changing command first checks that all the metadata nodes are in sync, it rejects from then on:

  • every citus_add_node, so a query router added later (or a new replica of an existing one) is never registered by Patroni;
  • distributed DDL.

The coordinator Patroni logs, on every heartbeat:

ERROR: Exception when executing query "SELECT pg_catalog.citus_add_node(%s, %s, %s, %s, 'default')", (('10.244.0.30', 5432, 1027, 'primary')): ObjectNotInPrerequisiteState('10.244.0.22:5432 is a metadata node, but is out of sync
HINT:  If the node is up, wait until metadata gets synced to it and try again.

Reproduction

  1. Create a citus SGShardedCluster and raise spec.coordinator.queryRouterClusters to 8.
  2. Lower it to 3 once the query routers are registered in pg_dist_node.
  3. The group of the removed query routers (e.g. 1032) is still in pg_dist_node, and the query router -router2 (group 1027), when recreated, is never added.

Expected result

The nodes of the workers and query routers whose group is no longer part of the sharded cluster (beyond workers.clusters, or beyond queryRouterIndexOffset + queryRouterClusters) are removed from pg_dist_node.

Resolution

  • Citus can only remove a primary node that can be reached, so the SGCluster of a removed worker or query router must not be scaled down to 0 instances while its group is registered in pg_dist_node. A new coordinator SGScript entry, citus-registered-groups (cron on updateNodeInterval, setValue), stores the groups registered in pg_dist_node in the coordinator SGCluster status. The SGShardedCluster does not scale down a removed worker or query router SGCluster whose group is among them, whatever the value of enableNodeAutoRemoval, and scales it down once its group is gone.
  • With the new SGShardedCluster.spec.configurations.citus.enableNodeAutoRemoval (disabled by default) the coordinator removes, with citus_remove_node, the nodes of the removed groups that hold no shard of a distributed table and can be reached. Without it they have to be removed manually on the coordinator primary (a worker holding shards has to be drained first).

Workaround

On the coordinator primary, for each stale query router that can no longer be reached (for example, one removed before this fix):

SELECT citus_disable_node('<nodename>', <nodeport>, synchronous => true);
SELECT citus_remove_node('<nodename>', <nodeport>);
Edited by Matteo Melli