Fix kube-ovn RAFT election

What does this MR do and why?

Fix for #4083 (closed) .

In order to fix the RAFT issue, we are forcing the node to leave the NB and SB process, deleting the node id in the RAFT cluster membership. With cluster leave, ovn controller also removes the node label ovn.kubernetes.io/allocated which forces the new ovn-controller pod to start the RAFT and OVSDB joining process after rolling upgrade.

The DB files are also cleared as "lock" files can interfere with the ovsdb negotiation process, they are repopulated after the rolling upgrade when the node gains RAFT membership.

Test coverage

Multiple tests ran locally :

[git:is-kovn-raftfix]root@vbmh3-feper:sylva-core# kubectl get pods -n kube-system | grep ovn
kube-ovn-cni-gdb5r                                     1/1     Running     0             16m
kube-ovn-cni-vcn4b                                     1/1     Running     0             73m
kube-ovn-controller-cf854498-5m5hz                     0/1     Pending     0             16m
kube-ovn-controller-cf854498-gr7k9                     1/1     Running     0             32m
kube-ovn-controller-cf854498-nhj89                     1/1     Running     1 (12m ago)   73m
kube-ovn-monitor-5d995f8b6d-c9qxd                      1/1     Running     0             16m
kube-ovn-pinger-l4gts                                  1/1     Running     0             73m
kube-ovn-pinger-qkbrz                                  1/1     Running     0             16m
ovn-central-8c649bb6-g29bq                             1/1     Running     1 (12m ago)   73m
ovn-central-8c649bb6-vqsfb                             1/1     Running     1 (12m ago)   32m
ovn-central-8c649bb6-zqqv9                             0/1     Pending     0             15m
ovs-ovn-b78pm                                          1/1     Running     0             73m
ovs-ovn-jv27j                                          1/1     Running     0             16m
[10:29 AM][git:is-kovn-raftfix]root@vbmh3-feper:sylva-core# kubectl get pods -n kube-system | grep ovn
kube-ovn-cni-46fsd                                     1/1     Running             0             3m2s
kube-ovn-cni-gdb5r                                     1/1     Running             0             47m
kube-ovn-cni-z9j4c                                     1/1     Running             0             23m
kube-ovn-controller-cf854498-5m5hz                     1/1     Running             0             46m
kube-ovn-controller-cf854498-gr7k9                     1/1     Running             0             63m
kube-ovn-controller-cf854498-hjb7f                     1/1     Running             0             22m
kube-ovn-monitor-5d995f8b6d-hgdg8                      1/1     Running             0             22m
kube-ovn-pinger-4q6x2                                  1/1     Running             0             23m
kube-ovn-pinger-6d9hl                                  1/1     Running             0             3m2s
kube-ovn-pinger-qkbrz                                  1/1     Running             0             47m
ovn-central-8c649bb6-fvqw2                             1/1     Running             0             22m
ovn-central-8c649bb6-vqsfb                             1/1     Running             1 (42m ago)   63m
ovn-central-8c649bb6-zqqv9                             1/1     Running             0             45m
ovs-ovn-6zwr4                                          1/1     Running             0             3m2s
ovs-ovn-g5ghg                                          1/1     Running             0             23m
ovs-ovn-jv27j                                          1/1     Running             0             47m

Successful LSP join for all new nodes:

[git:is-kovn-raftfix]root@vbmh3-feper:sylva-core# kubectl -n kube-system exec -it ovn-central-75fc9c8947-lxrpt -- ovn-nbctl show | grep -B2 -A5 "100.64.0.3\|join"
Defaulted container "ovn-central" out of: ovn-central, hostpath-init (init)
        type: router
        router-port: ovn-cluster-ovn-default
switch 7af0f98b-564e-4e2c-813b-1f9889592b9e (join)
    port node-management-cluster-server2
        addresses: ["02:9e:b4:f9:63:20 100.64.0.2"]
    port node-management-cluster-server1
        addresses: ["16:06:15:05:cb:c9 100.64.0.5"]
    port join-ovn-cluster
        type: router
        router-port: ovn-cluster-join
    port node-management-cluster-server3
        addresses: ["42:68:03:ef:24:f2 100.64.0.3"]
router e8f4de56-99b1-4637-8602-3a95d2620774 (ovn-cluster)
    port ovn-cluster-join
        mac: "d2:b0:36:60:a7:41"
        ipv6-lla: "fe80::d0b0:36ff:fe60:a741"
        networks: ["100.64.0.1/16"]
    port ovn-cluster-ovn-default
        mac: "2a:66:cd:57:a0:19"

CI configuration

Below you can choose test deployment variants to run in this MR's CI.

Click to open to CI configuration

Legend:

Icon Meaning Available values
☁️ Infra Provider capd, capo, capm3
🚀 Bootstrap Provider kubeadm (alias kadm), rke2, okd, ck8s
🐧 Node OS ubuntu, suse, na, leapmicro
🛠️ Deployment Options Deployment option list and description
🎬 Pipeline Scenarios Available scenario list and description
🟢 Enabled units Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type
🔴 Disabled units Any available units name, by default apply to management and workload cluster. Can be prefixed by mgmt: or wkld: to be applied only to a specific cluster type
🏗️ Target platform Can be used to select specific deployment environment Available platform list and description
⚡ Pipeline control autorun, manual or blocking. Can be used to override global config and start a deployment pipeline the required way
  • 🎬 preview ☁️ capd 🚀 kadm 🐧 ubuntu

  • 🎬 preview ☁️ capo 🚀 rke2 🐧 suse

  • 🎬 preview ☁️ capm3 🚀 rke2 🐧 ubuntu

  • ☁️ capd 🚀 kadm 🛠️ light-deploy 🐧 ubuntu

  • ☁️ capd 🚀 rke2 🛠️ light-deploy 🐧 suse

  • ☁️ capo 🚀 rke2 🐧 suse

  • ☁️ capo 🚀 rke2 🐧 leapmicro

  • ☁️ capo 🚀 kadm 🐧 ubuntu

  • ☁️ capo 🚀 kadm 🐧 ubuntu 🟢 neuvector,mgmt:harbor

  • ☁️ capo 🚀 rke2 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capo 🚀 kadm 🎬 wkld-k8s-upgrade 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🎬 rolling-update-no-wkld 🛠️ ha 🐧 suse

  • ☁️ capo 🚀 rke2 🎬 sylva-upgrade 🛠️ ha 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🎬 sylva-upgrade-from-1.6.x 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capo 🚀 rke2 🛠️ misc 🐧 ubuntu 🟢 mgmt:harbor 🔴 neuvector

  • ☁️ capo 🚀 rke2 🛠️ ha,misc,openbao🐧 suse

  • ☁️ capo 🚀 rke2 🐧 suse 🎬 upgrade-from-prev-tag

  • ☁️ capm3 🚀 rke2 🐧 suse

  • ☁️ capm3 🚀 kadm 🐧 ubuntu

  • ☁️ capm3 🚀 ck8s 🐧 ubuntu

  • ☁️ capm3 🚀 kadm 🎬 rolling-update-no-wkld 🛠️ ha,misc 🐧 ubuntu

  • ☁️ capm3 🚀 rke2 🎬 wkld-k8s-upgrade 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 kadm 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capm3 🚀 rke2 🎬 upgrade-from-prev-release-branch 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 rke2 🛠️ misc,ha 🐧 suse

  • ☁️ capm3 🚀 rke2 🎬 sylva-upgrade 🛠️ ha,misc 🐧 suse

  • ☁️ capm3 🚀 kadm 🎬 rolling-update 🛠️ ha 🐧 suse

  • ☁️ capm3 🚀 ck8s 🎬 rolling-update 🛠️ ha 🐧 ubuntu

  • ☁️ capm3 🚀 rke2|okd 🎬 no-update 🐧 ubuntu|na

  • ☁️ capm3 🚀 rke2 🐧 suse 🎬 upgrade-from-release-1.5

  • ☁️ capm3 🚀 rke2 🐧 suse 🎬 upgrade-to-main

  • ☁️ capm3 🚀 rke2 🐧 suse 🟢 kube-ovn,multus

  • ☁️ capo 🚀 rke2 🐧 suse 🟢 kube-ovn,multus

  • ☁️ capo 🚀 rke2 🎬 rolling-update 🛠️ ha 🐧 suse 🟢 kube-ovn,multus

  • ☁️ capm3 🚀 rke2 🎬 rolling-update 🛠️ ha 🐧 suse 🟢 kube-ovn,multus 🏗️ virt-leaseweb

Global config for deployment pipelines

  • autorun pipelines
  • allow failure on pipelines
  • record sylvactl events

Notes:

  • Enabling autorun will make deployment pipelines to be run automatically without human interaction
  • Disabling allow failure will make deployment pipelines mandatory for pipeline success.
  • if both autorun and allow failure are disabled, deployment pipelines will need manual triggering but will be blocking the pipeline

Be aware: after configuration change, pipeline is not triggered automatically. Please run it manually (by clicking the run pipeline button in Pipelines tab) or push new code.

Edited by Ionut Spanu

Merge request reports

Loading
Loading