Restart the primary without switchover only when hot standby sensitive parameters are decreased and add a configurable restart delay
Summary
Two related improvements to the rollout performed by ClusterStatefulSetWithPrimaryReconciliationHandler:
- Restart the Postgres instance of the primary Pod in place (that is, without performing a switchover) only when a hot standby sensitive parameter is being decreased. In any other case a switchover has to be performed.
- Add a configurable restart delay (
SGCluster.spec.pods.updateStrategy.restartDelay, defaultPT5M) enforced by the cluster controller, so that once a restart has been issued the next one has to wait at least such delay.
1. Restart the primary without switchover only when hot standby sensitive parameters are decreased
ClusterStatefulSetWithPrimaryReconciliationHandler.performRollout currently restarts the Postgres instance of the primary Pod, without performing any switchover, whenever
ClusterRolloutUtil.getPostgresRestartReasons(pod, patroniMembers).requiresRestart() is true:
if (foundPrimaryPod
.map(pod -> ClusterRolloutUtil.getPostgresRestartReasons(pod, patroniMembers)
.requiresRestart())
.orElse(false)) {
...
patroniCtl.restart(credentials.v1, credentials.v2,
foundPrimaryPod.get().getMetadata().getName());
return;
}This makes the primary suffer the downtime of a Postgres restart even when the change could have been applied with a switchover, that has a much lower impact on the availability of the cluster.
Restarting the primary in place is only required when any of the following parameters is decreased:
max_connectionsmax_prepared_transactionsmax_locks_per_transactionmax_wal_sendersmax_worker_processes
See https://www.postgresql.org/docs/current/hot-standby.html#HOT-STANDBY-ADMIN: a hot standby requires those parameters to be greater than or equal to the values used on the primary, therefore when they are decreased the primary has to apply the new (lower) values first, otherwise the replicas would refuse to continue the recovery and shut down. When those parameters are increased (or when any other parameter changes) the replicas have to be restarted first and the primary has to be switched over.
The parameters that are pending a change are exposed by patroni:
-
In the Pod
statusannotation (and in the REST API) through thepending_restart_reasonfield, for instance:{"pending_restart_reason":{"max_connections":{"old_value":"80","new_value":"79"}}} -
In
patronictl list -e --format jsonthrough thePending restart reasonfield, in a textual format with a parameter per line, for instance:max_connections: 80->79
Proposed changes
- Expose the pending restart reason in
PatroniMember(Pending restart reasonproperty), filling it also inPatroniCtlKubernetesInstancefrom thepending_restart_reasonfield of the Podstatusannotation, and provide a parsed representation of it. - Add a helper to
ClusterRolloutUtilthat returns whether any of the hot standby sensitive parameters listed above is being decreased for a given Pod. - In
performRollout, restart the Postgres instance of the primary Pod in place only when such helper returnstrue. In any other case continue with the restart of the Postgres instance of the replicas and, once no replica is pending a restart anymore, perform a switchover to the ready replica with the least lag. Fall back to an in place restart of the primary only if no switchover candidate is available. - Patroni only reports the reason of a pending restart since version 4. When it is not reported it is not possible to tell if a hot standby sensitive parameter is being decreased, therefore a switchover is performed.
- Document the behavior in the
restart,minorVersionUpgradeandsecurityUpgradesections of theSGDbOpsandSGShardedDbOpsCRDs and in theupdateStrategysection of theSGClusterandSGShardedClusterCRDs. - Add a check to the
dbops-restarte2e test that increasemax_connectionsand verify a switchover is performed, then decrease it and verify no switchover is performed. In both cases verify that the Pods are not re-created and that the Postgres instance of the primary is not restarted more than once.
2. Configurable restart delay
Add the field SGCluster.spec.pods.updateStrategy.restartDelay, an ISO-8601 duration with default PT5M, that applies to the restart operation controlled by the cluster controller: after a restart has been issued the next restart has to wait at least such delay.
This is enforced by the cluster controller by adding a timestamp to the value of the stackgres.io/patroni-operation Pod annotation (StackGresContext.PATRONI_OPERATION_KEY) after the restart has been performed. The annotation is not removed until the delay has passed, and since the operator waits for the annotation to be removed in order to consider the restart completed, no other restart can be issued in the meantime.
Proposed changes
- Add
restartDelaytoStackGresClusterUpdateStrategyand to theSGClusterandSGShardedClusterCRDs with defaultPT5M. - In
PatroniOperationReconciliator, after invoking the restart set arestartedtimestamp in thestackgres.io/patroni-operationannotation value instead of removing the annotation, and remove the annotation only oncerestartedplus the configured restart delay has passed.
Note that, since the operator waits for the annotation to be removed before issuing any other restart, the default of
PT5Mpaces a rollout at one Postgres instance restart every 5 minutes.