Kubelet升级是否需排空节点?Kubernetes跨版本升级咨询
Hey there, let’s tackle your Kubernetes upgrade questions step by step—this is a tricky scenario, but I’ve got practical insights to share based on real-world experience.
First, let’s clarify the best practice for node upgrades, since that’s where your container restart concern comes in.
Recommended Approach: Drain Nodes Before Upgrading Kubelet
The safest way to upgrade worker nodes is to drain them first, then upgrade the kubelet, and finally uncordon the node. Here’s the step-by-step for each node:
- Run
kubectl drain <node-name> --ignore-daemonsetsto gracefully evict pods from the node. This sends aSIGTERMto pods, waits for their termination grace period, and reschedules them to other healthy nodes. - Upgrade the kubelet package (using your OS package manager like
aptoryum) to v1.8.12. - Verify the kubelet is running correctly with
systemctl status kubelet. - Run
kubectl uncordon <node-name>to allow new pods to be scheduled on the node.
Drain vs. In-Place Upgrade: Key Differences
Let’s break down why draining is better than upgrading in-place:
- In-place upgrade: When you directly upgrade the kubelet on a running node, the kubelet will restart. Due to the spec hash change issue you mentioned, all containers on the node will abruptly restart. This risks:
- Unplanned service downtime, especially for pods with no replicas or critical workloads.
- No graceful termination—pods are killed immediately, which can lead to data loss or inconsistent application state.
- If the kubelet upgrade fails, the node goes offline with pods still running on it, making recovery harder.
- Drain first: This method ensures pods are gracefully moved to other nodes before any changes. Benefits include:
- Zero service downtime (assuming you have enough replicas and cluster capacity).
- Pods shut down properly, adhering to their termination grace periods.
- If the kubelet upgrade fails, the node is already cordoned, so no new pods are scheduled there—you can troubleshoot without impacting live traffic.
Short answer: Not recommended, and there are significant constraints that make this risky.
Key Limiting Factors
Kubernetes has strict version compatibility rules, and skipping two minor versions (1.7 → 1.10 skips 1.8 and 1.9) violates official best practices. Here’s why:
- API Version Compatibility: Kubernetes changes API versions between minor releases. For example,
extensions/v1beta1Deployments (used in v1.7) were deprecated in later versions and replaced withapps/v1. A direct jump to v1.10 could cause the control plane to fail to recognize old API objects, leading to broken workloads. - Kubelet-Apiserver Version Gap: The official rule is that kubelet versions can be at most one minor version behind the apiserver. If you upgrade the apiserver to v1.10 first, your v1.7 kubelets will be two versions behind and will fail to communicate with the apiserver—resulting in all nodes going offline.
- Etcd Data Format Changes: Each minor release may update how cluster state is stored in etcd. Skipping versions means the v1.10 control plane may not be able to parse the older v1.7 state data, leading to corruption or cluster failure.
- Unsupported Upgrade Path: Kubernetes does not test or support skipping two minor versions. If you run into issues, you won’t get official support, and troubleshooting will be far more complex.
Alternative to Minimize Restarts
If you want to reduce the number of container restart events, consider:
- Upgrading from v1.7.10 → v1.9.x first (skipping one minor version, which is sometimes allowed with caution), then to v1.10.3. Even this carries some risk, so test it in a staging cluster first.
- Schedule upgrades during low-traffic windows, and use
PodDisruptionBudgetsto ensure critical workloads always have enough replicas running during the upgrade.
At the end of the day, the safest path is to follow the official step-by-step upgrade flow: v1.7.10 → v1.8.12 → v1.9.x → v1.10.3. It may take more time, but it minimizes the risk of cluster outages or data loss.
内容的提问来源于stack exchange,提问作者Kun Li

