You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Container Service(ACS):如何安全重启Windows工作节点?

Hey there, let's tackle this issue with your Windows Kubernetes node not recovering after a reboot. I've run into similar scenarios with Windows nodes on Azure, so here's a step-by-step breakdown of what to check and fix:

Troubleshooting Windows Kubernetes Node Recovery After Reboot on Azure

First, let's align on your workflow: you drained the Windows worker node (state became Ready, SchedulingDisabled), shut it down via Azure Portal (state shifted to NotReady, SchedulingDisabled), then rebooted it—but it's stuck in an abnormal state. Here's how to resolve this:

1. Verify the Kubelet Service is Running on the Windows Node

Windows nodes often face issues where the Kubelet fails to auto-start after a reboot. To check this:

  • RDP into the problematic Windows VM (you can grab RDP credentials directly from the Azure Portal for the node)
  • Open Services (run services.msc from the command prompt or Start menu)
  • Locate the kubelet service:
    • If it's not running, start it manually
    • Set its startup type to Automatic to prevent this from recurring
  • For deeper insight, check the Kubelet logs for errors with this command:
    Get-Content C:\k\kubelet.log | Select-Object -Last 50
    
    Look for control plane connection issues or certificate errors—these are the most common culprits here.

2. Validate Network Connectivity Between Node and Control Plane

A reboot can reset network settings, blocking the node from communicating with the Kubernetes control plane:

  • On the Windows node, ping your control plane's API server endpoint (you can get this via kubectl cluster-info)
  • If ping fails, check the VM's Network Security Group (NSG) in Azure Portal: ensure ports 6443 (K8s API) and 10250 (Kubelet) are open inbound for the node
  • Also confirm the node has a valid IP address and can resolve the control plane's DNS name

3. Re-enable Scheduling and Check Node Status

Once the node is back online, you need to uncordon it to allow pod scheduling again. First, confirm the node is reachable:

  • Run these commands from your local machine (where kubectl is configured):
    kubectl uncordon <your-node-name>
    
  • Then check the node's current status:
    kubectl get nodes <your-node-name>
    
  • If it's still NotReady, get detailed troubleshooting info with:
    kubectl describe node <your-node-name>
    
    Look at the Conditions section—it will explicitly state why the node is marked NotReady (e.g., KubeletNotReady with a message about an unresponsive Kubelet)

4. Force the Node to Re-register with the Control Plane

If the node's registration is stuck, trigger a fresh registration:

  • On the Windows node, stop the Kubelet service first
  • Delete the node's kubeconfig and certificate files (typically located at C:\k\config and C:\var\lib\kubelet\pki)
  • Restart the Kubelet service—it will re-register with the control plane using new credentials
  • Wait 5-10 minutes, then run kubectl get nodes to check if it returns to Ready status

5. Check Azure VM Health

Sometimes the VM itself has underlying issues post-reboot:

  • Navigate to the problematic node's VM in Azure Portal
  • Check the Boot diagnostics tab to look for boot errors
  • Review Performance metrics to ensure CPU/memory usage is within normal ranges
  • If the VM is in a broken state, try re-deploying it via Azure Portal (this preserves the OS disk but restarts the VM on a new host)

内容的提问来源于stack exchange,提问作者HarshaP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 08:26:00