Deleting an SGCluster does not wait for its Pods to terminate, so a cluster recreated with the same name can deadlock

Summary

Deleting an SGCluster completes as soon as the resource itself is gone. The Pods of its StatefulSet are still running at that point, and they keep running for the whole termination grace period.

If an SGCluster with the same name and namespace is created in that window, the operator creates a fresh set of Patroni DCS endpoints for it — but the previous incarnation's Pods are still alive and their Patroni processes write into those brand-new endpoints. They take the leader lock and stamp the initialize key. When they are finally killed, the leader key goes away with them and initialize stays behind.

The new members then start with empty data directories against a DCS that says the cluster is already initialized but has no leader. None of them is allowed to bootstrap, so they all remain at waiting for leader to bootstrap. The cluster never starts and does not recover on its own.

Impact

  • A recreated SGCluster can deadlock permanently at startup. Getting out of it requires manually deleting the Patroni DCS endpoints of that cluster.
  • It is a race against the Pod termination grace period, so it happens intermittently.
  • Reachable by any flow that deletes an SGCluster and recreates it under the same name without waiting for the previous Pods to terminate — restore-in-place procedures, GitOps reconciliation that recreates a resource, a pipeline or operator retry after a failed deletion.
  • Everything else follows from this. Deleting an SGShardedCluster deletes the SGCluster resources it owns, so a sharded cluster recreated with the same name hits the same problem on its coordinator, worker and query router clusters.

What this is not

Stating it up front because it is the obvious first hypothesis and it is wrong: this is not the new cluster adopting leftover resources.

AbstractConciliator.evalReconciliationState (lines 105-134) already covers that case. A required resource that is deployed and owned by another resource makes the conciliator log a warning and return an empty ReconciliationResult until the leftover is gone, and it behaves as intended here — a recreated cluster gets a fresh SGCluster uid, fresh DCS endpoints and a fresh backup path.

The gap is one layer below: the conciliator reconciles resources, but Pods outlive the resources they belong to, and nothing prevents a Pod of a deleted SGCluster from writing to the DCS on its way out.

Proposed resolution

Add a finalizer on SGCluster that, on deletion, sets the instances to 0 and waits for the Pods generated by the StatefulSet to be garbage collected before the finalizer is released. Deletion of an SGCluster then only completes once no Pod of that cluster is left that could still write to the DCS, which also makes the derived SGShardedCluster case safe.

Deletions that must not wait

The finalizer only acts on what the Kubernetes garbage collector would delete anyway, so it can not terminate Pods the user asked to keep, and it can not make an SGCluster impossible to delete:

  • Orphan deletion: when the SGCluster is deleted with --cascade=orphan (propagationPolicy: Orphan, which sets the orphan finalizer), or its StatefulSet is not owned by the SGCluster anymore (the garbage collector already orphaned it), the finalizer is released without touching the StatefulSet or its Pods.
  • Ownership: only the Pods owned by the StatefulSet or by the SGCluster, and the ones without owner (a primary marked as non disruptible is released by the StatefulSet), are deleted and waited for.
  • Bounded wait: when all the remaining Pods are still terminating 2 minutes after their termination grace period has elapsed (e.g. their node is unreachable), the finalizer is released and a ClusterPodsTerminationTimeout warning event is sent. Pods are never force deleted.
  • No operator: the finalizer can only be removed by the operator. The uninstall guide documents that the SGClusters must be deleted before the operator, and how to remove the finalizer manually (kubectl patch sgcluster <name> --type json -p '[{"op":"remove","path":"/metadata/finalizers"}]').

Also part of the resolution

The e2e suite currently works around this. wait_cluster_resources_removed in stackgres-k8s/e2e/utils/cluster, called at the end of remove_cluster and remove_sharded_cluster, waits for the SGClusters, Pods, persistent volume claims and Patroni DCS endpoints of a release to disappear before the test recreates the cluster. That workaround should be removed once the finalizer is in place.

Note that the citus pg_dist_node wait added to wait_sharded_cluster in the same file is unrelated to this issue and should stay.

Edited by Matteo Melli