K3s Cluster Maintenance and Upgrades
Cluster Maintenance
Section titled “Cluster Maintenance”This guide covers routine maintenance, operating system patching, automated K3s upgrades, and volume backup procedures for bare-metal nodes.
Routine Cluster Health Checks
Section titled “Routine Cluster Health Checks”Perform routine health checks across cluster nodes, pods, and GitOps sync state:
kubectl get nodes -o widekubectl get pods -Akubectl get applications -n argocdConfirm that the local storage path /home/k3s-storage has ample free disk space:
ssh sudhanva@100.66.139.118 "df -h /home/k3s-storage"Node Maintenance and OS Patching
Section titled “Node Maintenance and OS Patching”When updating the host operating system (Ubuntu 26.04 LTS) or applying kernel updates, drain workloads cleanly before rebooting.
Step 1: Cordon the node
Section titled “Step 1: Cordon the node”Mark the node as unschedulable to prevent new pods from being assigned:
kubectl cordon legionStep 2: Drain active workloads
Section titled “Step 2: Drain active workloads”Evict running pods gracefully while honoring DaemonSets and local data volumes:
kubectl drain legion --ignore-daemonsets --delete-emptydir-data --forceStep 3: Apply host updates and reboot
Section titled “Step 3: Apply host updates and reboot”Log in to the host machine, run package updates, and reboot the node:
ssh sudhanva@100.66.139.118sudo apt update && sudo apt upgrade -ysudo rebootStep 4: Uncordon the node
Section titled “Step 4: Uncordon the node”After the node boots up and reports Ready, uncordon it to resume scheduling:
kubectl uncordon legionVerify that all workloads in media, argocd, vault, and monitoring return to Running state:
kubectl get pods -AAutomated K3s Upgrades with System Upgrade Controller
Section titled “Automated K3s Upgrades with System Upgrade Controller”The cluster utilizes Rancher System Upgrade Controller (infrastructure/system-upgrade-controller/) to execute automated, zero-downtime upgrades of K3s.
Step 1: Update the upgrade plan
Section titled “Step 1: Update the upgrade plan”Open infrastructure/system-upgrade-controller/k3s-upgrade-plan.yaml and update the target K3s version tag in spec.version:
apiVersion: upgrade.cattle.io/v1kind: Planmetadata: name: k3s-server namespace: system-upgradespec: concurrency: 1 nodeSelector: matchExpressions: - key: node.homelab/role operator: In values: - controlplane serviceAccountName: system-upgrade cordon: true drain: force: true ignoreDaemonSets: true deleteEmptydirData: true version: "v1.36.5+k3s1"Step 2: Commit and push changes
Section titled “Step 2: Commit and push changes”Commit the manifest change and push to master:
git add infrastructure/system-upgrade-controller/k3s-upgrade-plan.yamlgit commit -m "chore: bump k3s upgrade plan to v1.36.5+k3s1"git push origin masterStep 3: Monitor the upgrade job
Section titled “Step 3: Monitor the upgrade job”ArgoCD syncs the updated Plan. The System Upgrade Controller deploys a privileged upgrade job that drains the node, installs the specified K3s binary, restarts the service, and uncordons the node:
kubectl get pods -n system-upgrade -wkubectl get nodesStorage and Application Backups
Section titled “Storage and Application Backups”Because all persistent storage is backed by the native Local-Path Provisioner on /home/k3s-storage, disaster recovery is straightforward.
Local-Path Volume Backups
Section titled “Local-Path Volume Backups”Create snapshot backups of application data directories located at /home/k3s-storage/:
ssh sudhanva@100.66.139.118 "sudo tar -czvf /home/k3s-storage-backup-$(date +%F).tar.gz /home/k3s-storage"Vault Unseal Keys
Section titled “Vault Unseal Keys”Ensure your five Shamir unseal keys and root token in vault-keys.json are securely preserved in an offline password manager. If the Vault pod is rescheduled or restarted, unseal Vault using:
kubectl -n vault exec -it vault-0 -- vault operator unseal <key>