SGShardedCluster: query routers removed by decreasing queryRouterClusters are never removed from pg_dist_node
Summary
When spec.coordinator.queryRouterClusters of a citus SGShardedCluster is decreased, the query
routers that are removed stay registered in pg_dist_node forever. Neither Patroni (that only acts
on the Citus groups present in the DCS) nor StackGres ever call citus_remove_node for them.
The stale row is a metadata node (hasmetadata = t) that can not be reached, so Citus marks it
metadatasynced = f and, since every metadata changing command first checks that all the metadata
nodes are in sync, it rejects from then on:
- every
citus_add_node, so a query router added later (or a new replica of an existing one) is never registered by Patroni; - distributed DDL.
The coordinator Patroni logs, on every heartbeat:
ERROR: Exception when executing query "SELECT pg_catalog.citus_add_node(%s, %s, %s, %s, 'default')", (('10.244.0.30', 5432, 1027, 'primary')): ObjectNotInPrerequisiteState('10.244.0.22:5432 is a metadata node, but is out of sync
HINT: If the node is up, wait until metadata gets synced to it and try again.Reproduction
- Create a citus
SGShardedClusterand raisespec.coordinator.queryRouterClustersto 8. - Lower it to 3 once the query routers are registered in
pg_dist_node. - The group of the removed query routers (e.g.
1032) is still inpg_dist_node, and the query router-router2(group1027), when recreated, is never added.
Expected result
The nodes of the workers and query routers whose group is no longer part of the sharded cluster
(beyond workers.clusters, or beyond queryRouterIndexOffset + queryRouterClusters) are removed
from pg_dist_node.
Resolution
- Citus can only remove a primary node that can be reached, so the
SGClusterof a removed worker or query router must not be scaled down to 0 instances while its group is registered inpg_dist_node. A new coordinatorSGScriptentry,citus-registered-groups(cron onupdateNodeInterval,setValue), stores the groups registered inpg_dist_nodein the coordinatorSGClusterstatus. TheSGShardedClusterdoes not scale down a removed worker or query routerSGClusterwhose group is among them, whatever the value ofenableNodeAutoRemoval, and scales it down once its group is gone. - With the new
SGShardedCluster.spec.configurations.citus.enableNodeAutoRemoval(disabled by default) the coordinator removes, withcitus_remove_node, the nodes of the removed groups that hold no shard of a distributed table and can be reached. Without it they have to be removed manually on the coordinator primary (a worker holding shards has to be drained first).
Workaround
On the coordinator primary, for each stale query router that can no longer be reached (for example, one removed before this fix):
SELECT citus_disable_node('<nodename>', <nodeport>, synchronous => true);
SELECT citus_remove_node('<nodename>', <nodeport>);