SGShardedDbOps-ReconciliationLoop errors on already-Completed operations with a misleading 'non existent SGShardedCluster' message
Split from #3215 (issue 4), reported by @geass. Please refer to #3215 for the full original report.
Summary
Once a SGShardedDbOps reaches a terminal state, the SGShardedDbOps-ReconciliationLoop keeps
reconciling it and fails on every cycle with IllegalArgumentException: ... target non existent SGShardedCluster <name>, even though the SGShardedCluster exists and is healthy. The message is
misleading and the errors only stop when the completed SGShardedDbOps object is deleted.
Environment
- StackGres operator:
1.19.0(GA) SGShardedCluster(Citus), 1 coordinator + 4 workers
Current behaviour
ERROR [io.st.op.conciliation] (SGShardedDbOps-ReconciliationLoop) Reconciliation of SGShardedDbOps
citus-qa.qa-pending-restart-20260901 failed: java.lang.IllegalArgumentException: SGShardedDbOps
citus-qa.qa-pending-restart-20260901 target non existent SGShardedCluster example-cluster
at io.stackgres.operator.conciliation.shardeddbops.StackGresShardedDbOpsContext.lambda$getShardedCluster$0(...)
at java.base/java.util.Optional.orElseThrow(Optional.java:403)
...Repeats every reconcile cycle. The referenced SGShardedCluster existed and was healthy throughout.
Deleting the completed SGShardedDbOps stopped the errors immediately. No downstream effect was
observed beyond the log noise.
Analysis
The exception is not a lookup or caching failure — the cluster is deliberately not looked up:
ShardedDbOpsClusterContextAppender.appendContext(stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/conciliation/shardeddbops/context/ShardedDbOpsClusterContextAppender.java:38-41) short-circuits for a completed operation and setsfoundShardedCluster(Optional.empty())without querying the API.- Something downstream in resource generation then dereferences
StackGresShardedDbOpsContext.getShardedCluster()(stackgres-k8s/src/operator/src/main/java/io/stackgres/operator/conciliation/shardeddbops/StackGresShardedDbOpsContext.java:37-44), whoseorElseThrowproduces the "target non existent SGShardedCluster" message — a message that is simply wrong in this situation. ShardedDbOpsJobsGenerator(.../conciliation/factory/shardeddbops/ShardedDbOpsJobsGenerator.java:49) does filter out completed operations, so the caller that still dereferences the cluster needs to be identified; the full stack trace was truncated in the report (see "Missing information" below).- The non-sharded path has the identical shape (
DbOpsClusterContextAppender.java:55-58+StackGresDbOpsContext), so whatever is fixed here should be checked forSGDbOpstoo.
Expected behaviour
- A terminal (
Completed/failed)SGShardedDbOpsshould reconcile to a clean no-op, not to anERRORon every cycle. - If a context field is intentionally left empty for completed operations, dereferencing it must either be unreachable or fail with an accurate message — never claim a resource does not exist when it was never looked up.
Missing information (to request from the reporter)
- The full, untruncated stack trace of the
IllegalArgumentException, to pin down which resource generator still callsgetShardedCluster()for a completed operation. - The
opof the completedSGShardedDbOps(the object name suggests therestartfrom issue 3 of #3215) and its final.status.