Draft: in 2-step upgrade, control MachineDeployments k8s version separately

This MR is meant to address #3902 (closed) to reduce node rolling updates during Kubernetes 2-steps upgrades.

Since MD nodes can lag behind a few versions, we can do the following for a 2-step (example): when upgrading from 1.33 to 1.35, during the 1.33 to 1.34 upgrade it is sufficient to only upgrade the CP nodes, and leave MD Nodes in 1.33, the MD Nodes will then be upgraded to 1.35 at the same time as the CP nodes, on the second step.


⚠️ The attempt in this MR does not work well, because when attempting to let the MachineDeployments lag behing, by having them use a k8s version older than the one used for the CP still results in updating the MachineDeployments (for k8s patch version, for settings defined in sylva-units under .cluster, for sylva-capi-cluster version, for OS images).

The following illustrates the problem:

  • starting point: sylva 1.6, kubeadm mgmt cluster (k8s 1.33.9)
  • apply.sh to upgrade to main / sylva 1.7
  • step 1:
    • update CP to 1.34
    • keep MD in 1.33
      • ... but in main, 1.33 is 1.33.10 and different OS images and different s-c-c version
      • so MachineDeployment-related manifests are still updated
      • going through 1.33.10 would be a useless MD nodes rolling update

This approach still would roughly work, because:

  • for mgmt cluster:
    • rke2 suppports a direct n->n+2 jump (but this would not necessarily be the case in the future)
    • kubeadm has a pre-flight check that prevents creating MD nodes in a version different from CP nodes
  • for workload clusters:
    • sylva upgrades and k8s version upgrades should in practice be different operations, so we wouldn't enter this case
    • ... unless people try to do both in the same operation

Overall, this "compute a different version for MD nodes" idea does not work well and it's worth trying something else.

I started another MR, this time relying on MachineDeployment.spec.paused: true: !7862 (merged)


This MR:

  • does a minor variable name change (wrong name used in current version of the code)
  • extends the computation of the currently observed version to distinguish CP nodes and MD nodes
  • adds a computation of the current version (the next one to apply) distinct for MachineDeployments
  • upgrade s-c-c to use sylva-projects/sylva-elements/helm-charts/sylva-capi-cluster!982 (merged)
  • set cluster.machine_deployment_default.k8s_version to the version computed specifically for MachineDeployement nodes
  • adjusts cluster-machines-ready check so that during the step where MachineDeployments are purposefully left lagging behind, no check is done on them -- TBC ⚠️

💡 best reviewed one commit at a time

This MR depends on sylva-projects/sylva-elements/helm-charts/sylva-capi-cluster!982 (merged)

CI configuration

Below you can choose test deployment variants to run in this MR's CI.

Click to open to CI configuration

Legend:

Icon Meaning Available values
☁️ Infra Provider capd, capo, capm3
🚀 Bootstrap Provider kubeadm (alias kadm), rke2, okd, ck8s
🐧 Node OS ubuntu, suse, na, leapmicro
🛠️ Deployment Options Deployment option list and description
🎬 Pipeline Scenarios Available scenario list and description
🟢 Enabled units Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type
🔴 Disabled units Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type
🏗️ Target platform Can be used to select specific deployment environment Available platform list and description
Pipeline control autorun, manual or blocking. Can be used to override global config and start a deployment pipeline the required way
  • 🎬 preview ☁️ capd 🚀 kadm 🐧 ubuntu

  • 🎬 preview ☁️ capo 🚀 rke2 🐧 suse

  • 🎬 preview ☁️ capm3 🚀 rke2 🐧 ubuntu

  • ☁️ capd 🚀 kadm 🛠️ light-deploy 🐧 ubuntu

  • ☁️ capd 🚀 rke2 🛠️ light-deploy 🐧 suse

  • ☁️ capo 🚀 rke2 🐧 suse

  • ☁️ capo 🚀 rke2 🐧 leapmicro

  • ☁️ capo 🚀 kadm 🐧 ubuntu

  • ☁️ capo 🚀 kadm 🐧 ubuntu 🟢 neuvector,mgmt:harbor

  • ☁️ capo 🚀 rke2 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capo 🚀 kadm 🎬 sylva-upgrade 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🎬 rolling-update-no-wkld 🛠️ ha 🐧 suse

  • ☁️ capo 🚀 rke2 🎬 sylva-upgrade 🛠️ ha 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🎬 sylva-upgrade-from-1.6.x 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🛠️ misc 🐧 ubuntu 🟢 mgmt:harbor 🔴 neuvector

  • ☁️ capo 🚀 rke2 🛠️ ha,misc,openbao🐧 suse

  • ☁️ capo 🚀 rke2 🐧 suse 🎬 upgrade-from-prev-tag

  • ☁️ capm3 🚀 rke2 🐧 suse

  • ☁️ capm3 🚀 kadm 🐧 ubuntu

  • ☁️ capm3 🚀 ck8s 🐧 ubuntu

  • ☁️ capm3 🚀 kadm 🎬 rolling-update-no-wkld 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capm3 🚀 rke2 🎬 wkld-k8s-upgrade 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 kadm 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capm3 🚀 rke2 🎬 upgrade-from-prev-release-branch 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 rke2 🛠️ misc,ha 🐧 suse

  • ☁️ capm3 🚀 kadm 🎬 sylva-upgrade 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 kadm 🎬 rolling-update 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 ck8s 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capm3 🚀 rke2|okd 🎬 no-update 🐧 ubuntu|na

  • ☁️ capm3 🚀 rke2 🐧 suse 🎬 upgrade-from-release-1.5

  • ☁️ capm3 🚀 rke2 🐧 suse 🎬 upgrade-to-main

Global config for deployment pipelines

  • autorun pipelines
  • allow failure on pipelines
  • record sylvactl events

Notes:

  • Enabling autorun will make deployment pipelines to be run automatically without human interaction
  • Disabling allow failure will make deployment pipelines mandatory for pipeline success.
  • if both autorun and allow failure are disabled, deployment pipelines will need manual triggering but will be blocking the pipeline

Be aware: after configuration change, pipeline is not triggered automatically. Please run it manually (by clicking the run pipeline button in Pipelines tab) or push new code.

Edited by Thomas Morin

Merge request reports

Loading
Loading