Deleting an SGCluster does not wait for its Pods to terminate, so a cluster recreated with the same name can deadlock
Summary
Deleting an SGCluster completes as soon as the resource itself is gone. The Pods of its
StatefulSet are still running at that point, and they keep running for the whole termination
grace period.
If an SGCluster with the same name and namespace is created in that window, the operator
creates a fresh set of Patroni DCS endpoints for it — but the previous incarnation's Pods are
still alive and their Patroni processes write into those brand-new endpoints. They take the
leader lock and stamp the initialize key. When they are finally killed, the leader key goes
away with them and initialize stays behind.
The new members then start with empty data directories against a DCS that says the cluster is
already initialized but has no leader. None of them is allowed to bootstrap, so they all remain
at waiting for leader to bootstrap. The cluster never starts and does not recover on its own.
Impact
- A recreated
SGClustercan deadlock permanently at startup. Getting out of it requires manually deleting the Patroni DCS endpoints of that cluster. - It is a race against the Pod termination grace period, so it happens intermittently.
- Reachable by any flow that deletes an
SGClusterand recreates it under the same name without waiting for the previous Pods to terminate — restore-in-place procedures, GitOps reconciliation that recreates a resource, a pipeline or operator retry after a failed deletion. - Everything else follows from this. Deleting an
SGShardedClusterdeletes theSGClusterresources it owns, so a sharded cluster recreated with the same name hits the same problem on its coordinator, worker and query router clusters.
What this is not
Stating it up front because it is the obvious first hypothesis and it is wrong: this is not the new cluster adopting leftover resources.
AbstractConciliator.evalReconciliationState (lines 105-134) already covers that case. A
required resource that is deployed and owned by another resource makes the conciliator log a
warning and return an empty ReconciliationResult until the leftover is gone, and it behaves
as intended here — a recreated cluster gets a fresh SGCluster uid, fresh DCS endpoints and a
fresh backup path.
The gap is one layer below: the conciliator reconciles resources, but Pods outlive the
resources they belong to, and nothing prevents a Pod of a deleted SGCluster from writing to
the DCS on its way out.
Proposed resolution
Add a finalizer on SGCluster that, on deletion, sets the instances to 0 and waits for the
Pods generated by the StatefulSet to be garbage collected before the finalizer is released.
Deletion of an SGCluster then only completes once no Pod of that cluster is left that could
still write to the DCS, which also makes the derived SGShardedCluster case safe.
Deletions that must not wait
The finalizer only acts on what the Kubernetes garbage collector would delete anyway, so it can
not terminate Pods the user asked to keep, and it can not make an SGCluster impossible to
delete:
- Orphan deletion: when the
SGClusteris deleted with--cascade=orphan(propagationPolicy: Orphan, which sets theorphanfinalizer), or its StatefulSet is not owned by theSGClusteranymore (the garbage collector already orphaned it), the finalizer is released without touching the StatefulSet or its Pods. - Ownership: only the Pods owned by the StatefulSet or by the
SGCluster, and the ones without owner (a primary marked as non disruptible is released by the StatefulSet), are deleted and waited for. - Bounded wait: when all the remaining Pods are still terminating 2 minutes after their
termination grace period has elapsed (e.g. their node is unreachable), the finalizer is released
and a
ClusterPodsTerminationTimeoutwarning event is sent. Pods are never force deleted. - No operator: the finalizer can only be removed by the operator. The uninstall guide documents
that the
SGClusters must be deleted before the operator, and how to remove the finalizer manually (kubectl patch sgcluster <name> --type json -p '[{"op":"remove","path":"/metadata/finalizers"}]').
Also part of the resolution
The e2e suite currently works around this. wait_cluster_resources_removed in
stackgres-k8s/e2e/utils/cluster, called at the end of remove_cluster and
remove_sharded_cluster, waits for the SGClusters, Pods, persistent volume claims and Patroni
DCS endpoints of a release to disappear before the test recreates the cluster. That workaround
should be removed once the finalizer is in place.
Note that the citus pg_dist_node wait added to wait_sharded_cluster in the same file is
unrelated to this issue and should stay.