Skip to content

K3s Cluster Maintenance and Upgrades

This guide covers routine maintenance, operating system patching, automated K3s upgrades, and volume backup procedures for bare-metal nodes.

Perform routine health checks across cluster nodes, pods, and GitOps sync state:

Terminal window
kubectl get nodes -o wide
kubectl get pods -A
kubectl get applications -n argocd

Confirm that the local storage path /home/k3s-storage has ample free disk space:

Terminal window
ssh sudhanva@100.66.139.118 "df -h /home/k3s-storage"

When updating the host operating system (Ubuntu 26.04 LTS) or applying kernel updates, drain workloads cleanly before rebooting.

Mark the node as unschedulable to prevent new pods from being assigned:

Terminal window
kubectl cordon legion

Evict running pods gracefully while honoring DaemonSets and local data volumes:

Terminal window
kubectl drain legion --ignore-daemonsets --delete-emptydir-data --force

Log in to the host machine, run package updates, and reboot the node:

Terminal window
ssh sudhanva@100.66.139.118
sudo apt update && sudo apt upgrade -y
sudo reboot

After the node boots up and reports Ready, uncordon it to resume scheduling:

Terminal window
kubectl uncordon legion

Verify that all workloads in media, argocd, vault, and monitoring return to Running state:

Terminal window
kubectl get pods -A

Automated K3s Upgrades with System Upgrade Controller

Section titled “Automated K3s Upgrades with System Upgrade Controller”

The cluster utilizes Rancher System Upgrade Controller (infrastructure/system-upgrade-controller/) to execute automated, zero-downtime upgrades of K3s.

Open infrastructure/system-upgrade-controller/k3s-upgrade-plan.yaml and update the target K3s version tag in spec.version:

apiVersion: upgrade.cattle.io/v1
kind: Plan
metadata:
name: k3s-server
namespace: system-upgrade
spec:
concurrency: 1
nodeSelector:
matchExpressions:
- key: node.homelab/role
operator: In
values:
- controlplane
serviceAccountName: system-upgrade
cordon: true
drain:
force: true
ignoreDaemonSets: true
deleteEmptydirData: true
version: "v1.36.5+k3s1"

Commit the manifest change and push to master:

Terminal window
git add infrastructure/system-upgrade-controller/k3s-upgrade-plan.yaml
git commit -m "chore: bump k3s upgrade plan to v1.36.5+k3s1"
git push origin master

ArgoCD syncs the updated Plan. The System Upgrade Controller deploys a privileged upgrade job that drains the node, installs the specified K3s binary, restarts the service, and uncordons the node:

Terminal window
kubectl get pods -n system-upgrade -w
kubectl get nodes

Because all persistent storage is backed by the native Local-Path Provisioner on /home/k3s-storage, disaster recovery is straightforward.

Create snapshot backups of application data directories located at /home/k3s-storage/:

Terminal window
ssh sudhanva@100.66.139.118 "sudo tar -czvf /home/k3s-storage-backup-$(date +%F).tar.gz /home/k3s-storage"

Ensure your five Shamir unseal keys and root token in vault-keys.json are securely preserved in an offline password manager. If the Vault pod is rescheduled or restarted, unseal Vault using:

Terminal window
kubectl -n vault exec -it vault-0 -- vault operator unseal <key>