SGShardedCluster PendingUpgrade condition is not aggregated from children and gives no actionable information

Split from #3215 (issue 2), reported by @geass. Please refer to #3215 for the full original report.

Summary

The PendingUpgrade condition on SGShardedCluster stayed True with a lastTransitionTime from five days earlier, through a full cluster restart and a confirmed minor version upgrade, while the equivalent condition on the child SGClusters was current. The condition gives the user no way to know what is actually pending or how to clear it.

Environment

  • StackGres operator: 1.19.0-rc5 upgraded to 1.19.0 (GA)
  • SGShardedCluster (Citus), 1 coordinator + 4 workers, single instance each

Current behaviour

On the SGShardedCluster:

Last Transition Time:  2026-08-17T06:28:48.556178Z
Reason:                ShardedClusterRequiresUpgrade
Status:                True
Type:                  PendingUpgrade

Unchanged through:

  • a full cluster restart (SGShardedDbOps op: restart, both ReducedImpact and InPlace, both reaching Completed)
  • a real, verified minor version upgrade (SGShardedDbOps op: minorVersionUpgrade, 17.9 -> 17.10, confirmed with SHOW server_version on every instance)

PendingRestart on the same object did get a fresh timestamp after the restart operations, so the condition-writing path itself works.

Analysis

ShardedClusterStatusManager.isPendingUpgrade (stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/conciliation/shardedcluster/ShardedClusterStatusManager.java:96-109) decides this condition exclusively from the stackgres.io/operator-version annotation on the SGShardedCluster itself: it is True whenever that annotation differs from the running operator version. It does not aggregate the child SGCluster conditions (unlike isPendingRestart right above it, which does), and it has nothing to do with Postgres versions.

Consequences:

  1. The condition is only cleared when the operator-version annotation is bumped, i.e. by a securityUpgrade operation. restart and minorVersionUpgrade legitimately do not clear it — so the observed "stale" state is the operator asking for a security upgrade that was never run, but nothing in the condition says so.
  2. lastTransitionTime not moving while status stays True is correct per Kubernetes condition conventions; the actual problem is the missing/unactionable information, not the timestamp.
  3. The comparison in the original report against the per-SGCluster message ("PostgreSQL 17.10 is the latest minor version for major 17. A newer major version 18.4 is available.") is comparing two different conditions: that message belongs to ComponentsUpdated, not to PendingUpgrade (stackgres-k8s/src/common/src/main/java/io/stackgres/common/crd/sgcluster/ClusterStatusCondition.java:22-30). The per-SGCluster PendingUpgrade (ClusterRequiresUpgrade) is also operator-version driven.

There is still a real gap: at the SGCluster level the upgrade need is derived from the actual cluster state (ClusterRolloutUtil.getClusterRestartReasons, reason UPGRADE), while at the SGShardedCluster level it is a single annotation check that ignores the children entirely. A sharded cluster whose children require an upgrade but whose own annotation happens to be current would report PendingUpgrade=False.

Expected behaviour

  • SGShardedCluster.PendingUpgrade should reflect the state of the sharded cluster and its children, consistently with how PendingRestart is aggregated.
  • The condition should carry a message stating what is pending and what clears it (a SGShardedDbOps op: securityUpgrade), so it is actionable instead of looking stuck.
  • Document the distinction between PendingUpgrade (operator version / security upgrade) and ComponentsUpdated (Postgres minor/major and extension versions) — the names invite exactly the confusion seen in this report.