SGShardedCluster reconciler: recurring ClusterUpdated churn, Services and Endpoints re-created every cycle

Split from #3215 (issue 1), reported by @geass. Please refer to #3215 for the full original report.

Summary

The SGShardedCluster reconciler emits a ClusterUpdated event roughly every 1-2 seconds, continuously, recreating/patching the same set of resources on every cycle. Observed on 1.19.0-rc5 and still present after upgrading to 1.19.0 GA.

Environment

  • StackGres operator: 1.19.0-rc5 and 1.19.0 (GA)
  • SGShardedCluster (Citus), 1 coordinator + 4 workers, single instance each (no HA replicas)
  • Kubernetes: cloud-managed

Current behaviour

Normal  ClusterUpdated  2s (x641337 over 14d)  stackgres  SGShardedCluster citus-qa.example-cluster
  updated: +Service:example-cluster-routers, +Service:example-cluster-workers,
  +Endpoints:example-cluster-workers, +Endpoints:example-cluster-routers and other 7
  resources where created, 3 resources where deleted, 168 resources where patched

The counter (x641337 over 14d) is the event's own age tracking, so this had been running for 14+ days at that rate. Note the shape of the summary: on every single cycle the same Services and Endpoints are reported as created, other resources as deleted, and ~168 as patched — a create/delete/patch loop that never converges.

After the 1.19.0-rc5 -> 1.19.0 upgrade the per-SGCluster churn stopped (consistent with the "wire in the skipUpdate conciliator hook" and pod-label-propagation changes in 1.19.0), but the SGShardedCluster-level churn continued unchanged, same rate, same pattern.

Impact

No data or pod-stability impact was observed by the reporter (two full cluster restarts and a minor version upgrade ran while this was firing, with no traceable disruption). But it is continuous, wasteful API server load, and the event/log noise makes genuine problems much harder to spot.

Analysis / starting points

  • The reporter suspected metadata.creationTimestamp / metadata.generation being part of the desired-vs-live diff, based on the per-SGCluster event text (StatefulSet:example-coord (+/metadata/creationTimestamp -> ...), (+/metadata/generation -> 2)). Worth noting when triaging: that text comes from PatchResumer (stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/common/PatchResumer.java:69-71), which renders a raw JsonDiff between the required and the deployed object, so server-managed fields will always appear there. The event text alone therefore does not prove what triggered the patch — but the fact that the same resources are created on every cycle does show a real, non-converging loop.
  • ShardedClusterConciliator (stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/conciliation/shardedcluster/ShardedClusterConciliator.java:53) only overrides skipUpdate for child SGClusters during a major version upgrade; everything else goes through the generic AbstractConciliator path.
  • The churning resources are exactly the ones produced by ShardedClusterServices and ShardedClusterEndpoints (stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/conciliation/factory/shardedcluster/ShardedClusterEndpoints.java). The generated Endpoints copy subsets from the coordinator/worker primary Endpoints into a new object; that is a plausible source of a permanent required-vs-deployed mismatch (and of a fight with the Kubernetes endpoints controller), so it is the first thing to check.
  • Related history: #2114 (closed) reported both ClusterBootstrapCompleted repetition and ClusterUpdated spam; !1335 (merged) (closing #1935 (closed) / #2226 (closed)) addressed the bootstrap-event part. The ClusterUpdated symptom described there appears never to have been fixed for the sharded path.

Expected behaviour

A steady-state SGShardedCluster with no spec or environment changes should reconcile to a no-op: no ClusterUpdated events, no repeated create/delete/patch of the same Services and Endpoints.

Suggested acceptance criteria

  • An e2e/integration check that a settled SGShardedCluster produces no ClusterUpdated event over a number of reconciliation cycles.