SGShardedCluster reconciliation reverts status changes made during a cycle, leaving status.dbOps set forever
Summary
ShardedClusterReconciliator.onPostReconciliation replaces the whole .status of the
SGShardedCluster with the copy that was loaded at the beginning of the reconciliation cycle:
clusterWriter.update(config, (currentShardedCluster) -> {
if (config.getStatus() != null) {
currentShardedCluster.setStatus(config.getStatus());
}
});AbstractCustomResourceWriter.update(resource, setter) re-reads the resource before applying the
setter, so metadata and resourceVersion are current, but the status is then overwritten wholesale
with the stale snapshot. Any change made to the status by another actor while the cycle was running
is silently reverted.
SGShardedCluster.status.dbOps is written and removed by the SGShardedDbOps job
(run-sharded-major-version-upgrade.sh sets it when the major version upgrade starts and removes it
when every child SGDbOps has completed). When the removal lands inside an in-flight reconciliation
cycle, the operator puts status.dbOps back. Nothing removes it afterwards: the job has already
exited successfully and the SGShardedDbOps is OperationCompleted, so the field stays set forever.
ClusterReconciliator.onPostReconciliation has carried the guard for exactly this since
430a56446f: it copies dbOps (together with os, arch, podStatuses and managedSql) from the
freshly read resource into the status it is about to write. The SGShardedCluster counterpart never
got it.
dbOps is the only field of StackGresShardedClusterStatus written from outside the reconciliation
cycle. The spec is not affected: AbstractReconciliator already guards it with a resourceVersion
equality check.
Impact
- While
status.dbOps.majorVersionUpgradeis set,ShardedClusterConciliator.skipUpdateandskipDeletionstop reconciling the childSGClusters. Astatus.dbOpsthat is never removed therefore freezes the children of theSGShardedClusterindefinitely: no configuration change, no scaling and no removal of obsolete resources is applied to them any more. - A subsequent
SGShardedDbOpsof the same kind findsstatus.dbOpsalready set and reuses itssourcePostgresVersioninstead of the real one. - The window is the duration of one reconciliation cycle, so it is intermittent. It is much easier
to hit than it looks, because after a major version upgrade the
SGShardedClusterreconciliation never converges: the pre-upgrade defaultSGPostgresConfigs<cluster>-{coordinator,worker N}-<old major>are still deployed and owned but are no longer required,AbstractConciliatorreports them as deletions on every cycle andFireAndForgetReconciliationHandlerdoes not delete them. The result is a continuous stream of reconciliation cycles, each one ending in a status write. That non-convergence is a separate defect of the same shape as #3219 and deserves its own fix.
Proposed resolution
Preserve the externally owned part of the status in
ShardedClusterReconciliator.onPostReconciliation, the way ClusterReconciliator does: read
status.dbOps from the resource that CustomResourceWriter.update has just fetched and carry it
over to the status being written, instead of overwriting it with the value from the start of the
cycle.
The e2e test sharded-dbops-major-version-upgrade already asserts the expected behaviour
("SGShardedCluster dbOps status was not cleaned up") and fails intermittently because of this.