Restart the primary without switchover only when hot standby sensitive parameters are decreased and add a configurable restart delay

Summary

Two related improvements to the rollout performed by ClusterStatefulSetWithPrimaryReconciliationHandler:

  1. Restart the Postgres instance of the primary Pod in place (that is, without performing a switchover) only when a hot standby sensitive parameter is being decreased. In any other case a switchover has to be performed.
  2. Add a configurable restart delay (SGCluster.spec.pods.updateStrategy.restartDelay, default PT5M) enforced by the cluster controller, so that once a restart has been issued the next one has to wait at least such delay.

1. Restart the primary without switchover only when hot standby sensitive parameters are decreased

ClusterStatefulSetWithPrimaryReconciliationHandler.performRollout currently restarts the Postgres instance of the primary Pod, without performing any switchover, whenever ClusterRolloutUtil.getPostgresRestartReasons(pod, patroniMembers).requiresRestart() is true:

if (foundPrimaryPod
    .map(pod -> ClusterRolloutUtil.getPostgresRestartReasons(pod, patroniMembers)
        .requiresRestart())
    .orElse(false)) {
  ...
  patroniCtl.restart(credentials.v1, credentials.v2,
      foundPrimaryPod.get().getMetadata().getName());
  return;
}

This makes the primary suffer the downtime of a Postgres restart even when the change could have been applied with a switchover, that has a much lower impact on the availability of the cluster.

Restarting the primary in place is only required when any of the following parameters is decreased:

  • max_connections
  • max_prepared_transactions
  • max_locks_per_transaction
  • max_wal_senders
  • max_worker_processes

See https://www.postgresql.org/docs/current/hot-standby.html#HOT-STANDBY-ADMIN: a hot standby requires those parameters to be greater than or equal to the values used on the primary, therefore when they are decreased the primary has to apply the new (lower) values first, otherwise the replicas would refuse to continue the recovery and shut down. When those parameters are increased (or when any other parameter changes) the replicas have to be restarted first and the primary has to be switched over.

The parameters that are pending a change are exposed by patroni:

  • In the Pod status annotation (and in the REST API) through the pending_restart_reason field, for instance:

    {"pending_restart_reason":{"max_connections":{"old_value":"80","new_value":"79"}}}
  • In patronictl list -e --format json through the Pending restart reason field, in a textual format with a parameter per line, for instance:

    max_connections: 80->79

Proposed changes

  • Expose the pending restart reason in PatroniMember (Pending restart reason property), filling it also in PatroniCtlKubernetesInstance from the pending_restart_reason field of the Pod status annotation, and provide a parsed representation of it.
  • Add a helper to ClusterRolloutUtil that returns whether any of the hot standby sensitive parameters listed above is being decreased for a given Pod.
  • In performRollout, restart the Postgres instance of the primary Pod in place only when such helper returns true. In any other case continue with the restart of the Postgres instance of the replicas and, once no replica is pending a restart anymore, perform a switchover to the ready replica with the least lag. Fall back to an in place restart of the primary only if no switchover candidate is available.
  • Patroni only reports the reason of a pending restart since version 4. When it is not reported it is not possible to tell if a hot standby sensitive parameter is being decreased, therefore a switchover is performed.
  • Document the behavior in the restart, minorVersionUpgrade and securityUpgrade sections of the SGDbOps and SGShardedDbOps CRDs and in the updateStrategy section of the SGCluster and SGShardedCluster CRDs.
  • Add a check to the dbops-restart e2e test that increase max_connections and verify a switchover is performed, then decrease it and verify no switchover is performed. In both cases verify that the Pods are not re-created and that the Postgres instance of the primary is not restarted more than once.

2. Configurable restart delay

Add the field SGCluster.spec.pods.updateStrategy.restartDelay, an ISO-8601 duration with default PT5M, that applies to the restart operation controlled by the cluster controller: after a restart has been issued the next restart has to wait at least such delay.

This is enforced by the cluster controller by adding a timestamp to the value of the stackgres.io/patroni-operation Pod annotation (StackGresContext.PATRONI_OPERATION_KEY) after the restart has been performed. The annotation is not removed until the delay has passed, and since the operator waits for the annotation to be removed in order to consider the restart completed, no other restart can be issued in the meantime.

Proposed changes

  • Add restartDelay to StackGresClusterUpdateStrategy and to the SGCluster and SGShardedCluster CRDs with default PT5M.
  • In PatroniOperationReconciliator, after invoking the restart set a restarted timestamp in the stackgres.io/patroni-operation annotation value instead of removing the annotation, and remove the annotation only once restarted plus the configured restart delay has passed.

Note that, since the operator waits for the annotation to be removed before issuing any other restart, the default of PT5M paces a rollout at one Postgres instance restart every 5 minutes.

Edited by Matteo Melli