Azure Container Service(ACS):如何安全重启Windows工作节点?
Hey there, let's tackle this issue with your Windows Kubernetes node not recovering after a reboot. I've run into similar scenarios with Windows nodes on Azure, so here's a step-by-step breakdown of what to check and fix:
First, let's align on your workflow: you drained the Windows worker node (state became Ready, SchedulingDisabled), shut it down via Azure Portal (state shifted to NotReady, SchedulingDisabled), then rebooted it—but it's stuck in an abnormal state. Here's how to resolve this:
1. Verify the Kubelet Service is Running on the Windows Node
Windows nodes often face issues where the Kubelet fails to auto-start after a reboot. To check this:
- RDP into the problematic Windows VM (you can grab RDP credentials directly from the Azure Portal for the node)
- Open Services (run
services.mscfrom the command prompt or Start menu) - Locate the
kubeletservice:- If it's not running, start it manually
- Set its startup type to Automatic to prevent this from recurring
- For deeper insight, check the Kubelet logs for errors with this command:
Look for control plane connection issues or certificate errors—these are the most common culprits here.Get-Content C:\k\kubelet.log | Select-Object -Last 50
2. Validate Network Connectivity Between Node and Control Plane
A reboot can reset network settings, blocking the node from communicating with the Kubernetes control plane:
- On the Windows node, ping your control plane's API server endpoint (you can get this via
kubectl cluster-info) - If ping fails, check the VM's Network Security Group (NSG) in Azure Portal: ensure ports 6443 (K8s API) and 10250 (Kubelet) are open inbound for the node
- Also confirm the node has a valid IP address and can resolve the control plane's DNS name
3. Re-enable Scheduling and Check Node Status
Once the node is back online, you need to uncordon it to allow pod scheduling again. First, confirm the node is reachable:
- Run these commands from your local machine (where
kubectlis configured):kubectl uncordon <your-node-name> - Then check the node's current status:
kubectl get nodes <your-node-name> - If it's still
NotReady, get detailed troubleshooting info with:
Look at the Conditions section—it will explicitly state why the node is marked NotReady (e.g.,kubectl describe node <your-node-name>KubeletNotReadywith a message about an unresponsive Kubelet)
4. Force the Node to Re-register with the Control Plane
If the node's registration is stuck, trigger a fresh registration:
- On the Windows node, stop the Kubelet service first
- Delete the node's kubeconfig and certificate files (typically located at
C:\k\configandC:\var\lib\kubelet\pki) - Restart the Kubelet service—it will re-register with the control plane using new credentials
- Wait 5-10 minutes, then run
kubectl get nodesto check if it returns toReadystatus
5. Check Azure VM Health
Sometimes the VM itself has underlying issues post-reboot:
- Navigate to the problematic node's VM in Azure Portal
- Check the Boot diagnostics tab to look for boot errors
- Review Performance metrics to ensure CPU/memory usage is within normal ranges
- If the VM is in a broken state, try re-deploying it via Azure Portal (this preserves the OS disk but restarts the VM on a new host)
内容的提问来源于stack exchange,提问作者HarshaP

