workload-cluster: robustness improvement for dummy units
Context: we've seen platforms that for historical reasons still were using units.sriov.enabled: true in workload clusters; this unit is a dummy unit (producing no resource), but workload context defaults give it a spec.kubeConfig to point to the workload cluster... as a result if the unit reconciles before the kubeconfig secret is created, things break.
We don't encounter this problem in CI because in workload-cluster.values.yaml the dummy units we use inherit from kustomization-deployed-on-mgmt-cluster (which clears the kubeConfig). A possibility could have been to fix the sriov unit to also have it inherit from kustomization-deployed-on-mgmt-cluster, or clear its kubeConfig. This is what we'll do to improve release branches (see #4392 (closed)).
But in fact, since dummy units never need a spec.kubeConfig (because they don't create resources), a cleaner solution is to have the dummy unit template clear the kubeConfig upfront. This is point (A).
... later edit:
- point (B): as pointed out below by @ishitamittal.2301, this is incorrect since we have some units inheriting from dummy which create resources
- for these the cleanest is to stop using the
dummyunit template
- for these the cleanest is to stop using the
- point (C): I later realized that another use-case for dummy unit is healthchecks done in the remote cluster: these need the kubeConfig as well
What this MR is doing:
- bring this fix (well, robustness improvement) which avoids having an irrelevant kubeConfig on dummy units
- do a related code simplification (cluster-ready unit does not anymore need to inherit from
kustomization-deployed-on-mgmt-cluster, so theunit_templatesinherited from values.yaml does not need to be overloaded) - do a minor cleanup: cluster-reachable dummy unit
unit_templatesstatement can be removed since it only repeats the one inherited from values.yaml (this was already true before this MR) - to address point (B), a separate commit adjusts some
xxx-tlsunits in workload-cluster.values.yaml to avoid using thedummyunit template - to address point (C), a separate commit introduces a
dummy-for-in-cluster-healthchecks, dedicated to units that today inherit fromdummyto do healthcheck in the remote cluster -- this solves point (C) and also simplifies the code a bit by avoiding the repeat of the kubeConfig (which these units had to point to the remote cluster not only in workload context, but also in bootstrap context)
CI configuration
Below you can choose test deployment variants to run in this MR's CI.
Click to open to CI configuration
Legend:
| Icon | Meaning | Available values |
|---|---|---|
| Infra Provider | capd, capo, capm3 |
|
| Bootstrap Provider | kubeadm (alias kadm), rke2, okd, ck8s |
|
| Node OS | ubuntu, suse, na, leapmicro |
|
| Deployment Options | Deployment option list and description | |
| Pipeline Scenarios | Available scenario list and description | |
| Enabled units | Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type |
|
| Disabled units | Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type |
|
| Target platform | Can be used to select specific deployment environment Available platform list and description | |
| Pipeline control | autorun, manual or blocking. Can be used to override global config and start a deployment pipeline the required way |
-
🎬 preview☁️ capd🚀 kadm🐧 ubuntu -
🎬 preview☁️ capo🚀 rke2🐧 suse -
🎬 preview☁️ capm3🚀 rke2🐧 ubuntu -
☁️ capd🚀 kadm🛠️ light-deploy🐧 ubuntu -
☁️ capd🚀 rke2🛠️ light-deploy🐧 suse -
☁️ capo🚀 rke2🐧 suse -
☁️ capo🚀 rke2🐧 leapmicro -
☁️ capo🚀 kadm🐧 ubuntu -
☁️ capo🚀 kadm🐧 ubuntu🟢 neuvector,mgmt:harbor -
☁️ capo🚀 rke2🎬 rolling-update🛠️ ha🐧 ubuntu -
☁️ capo🚀 kadm🎬 wkld-k8s-upgrade🐧 ubuntu -
☁️ capo🚀 rke2🎬 rolling-update-no-wkld🛠️ ha🐧 suse -
☁️ capo🚀 rke2🎬 sylva-upgrade🛠️ ha🐧 ubuntu -
☁️ capo🚀 rke2🎬 sylva-upgrade-from-1.6.x🛠️ ha,misc🐧 ubuntu -
☁️ capo🚀 rke2🛠️ ha,misc🐧 ubuntu -
☁️ capo🚀 rke2🛠️ misc🐧 ubuntu🟢 mgmt:harbor🔴 neuvector -
☁️ capo🚀 rke2🛠️ ha,misc,openbao🐧 suse -
☁️ capo🚀 rke2🐧 suse🎬 upgrade-from-prev-tag -
☁️ capm3🚀 rke2🐧 suse -
☁️ capm3🚀 kadm🐧 ubuntu -
☁️ capm3🚀 ck8s🐧 ubuntu -
☁️ capm3🚀 kadm🎬 rolling-update-no-wkld🛠️ ha,misc🐧 ubuntu -
☁️ capm3🚀 rke2🎬 wkld-k8s-upgrade🛠️ ha🐧 suse -
☁️ capm3🚀 kadm🎬 rolling-update🛠️ ha🐧 ubuntu -
☁️ capm3🚀 rke2🎬 upgrade-from-prev-release-branch🛠️ ha🐧 suse -
☁️ capm3🚀 rke2🛠️ misc,ha🐧 suse -
☁️ capm3🚀 rke2🎬 sylva-upgrade🛠️ ha,misc🐧 suse -
☁️ capm3🚀 kadm🎬 rolling-update🛠️ ha🐧 suse -
☁️ capm3🚀 ck8s🎬 rolling-update🛠️ ha🐧 ubuntu -
☁️ capm3🚀 rke2|okd🎬 no-update🐧 ubuntu|na -
☁️ capm3🚀 rke2🐧 suse🎬 upgrade-from-release-1.5 -
☁️ capm3🚀 rke2🐧 suse🎬 upgrade-to-main
Global config for deployment pipelines
- autorun pipelines
- allow failure on pipelines
- record sylvactl events
Notes:
- Enabling
autorunwill make deployment pipelines to be run automatically without human interaction - Disabling
allow failurewill make deployment pipelines mandatory for pipeline success. - if both
autorunandallow failureare disabled, deployment pipelines will need manual triggering but will be blocking the pipeline
Be aware: after configuration change, pipeline is not triggered automatically.
Please run it manually (by clicking the run pipeline button in Pipelines tab) or push new code.