如何在不删除Azure Kubernetes集群的情况下停止(取消分配)节点?
Got it, let’s work through this problem together. The reason you can only scale down to 1 node is almost certainly because you’re dealing with the system node pool—Azure requires these to have at least one running node to host critical Kubernetes system pods like kube-proxy, coredns, and the controller manager. But we’ve got ways to either eliminate weekend costs entirely or slash them to near-zero, depending on how your cluster is set up.
Option 1: You have separate user node pools (recommended setup)
If you’ve deployed your application workloads on dedicated user node pools (not the system pool), you can scale those down to 0 completely, and optimize the system pool to minimize its cost.
Step-by-step actions:
- Identify your node pools first
Run this Azure CLI command to list all node pools and spot which one is the system pool (look for theSystem: Yesentry):az aks nodepool list --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME -o table - Scale user node pools to 0
For every user node pool you have, run this command to deallocate all its VMs (you won’t be charged for these over the weekend):az aks nodepool scale --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME --name USER_NODE_POOL_NAME --node-count 0 - Optimize the system node pool (for extra savings)
If you want to cut costs even more, resize the system pool’s VM to a cheaper SKU (likeStandard_B1sorStandard_A1_v2) before locking it to 1 node:
Wait for the upgrade to finish, then confirm it’s scaled to 1 node:az aks nodepool upgrade --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME --name SYSTEM_NODE_POOL_NAME --node-vm-size Standard_B1s --no-waitaz aks nodepool scale --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME --name SYSTEM_NODE_POOL_NAME --node-count 1
Option 2: You only have a system node pool (no user pools)
If all your workloads are running on the system node pool, you can’t scale it to 0—but you can stop and deallocate the underlying VMs to avoid VM compute charges. Note: This will make your cluster unavailable over the weekend, so only do this if you don’t need access to it during that time.
Step-by-step actions:
- Find the VM scale set tied to your system node pool
First, grab the scale set ID with this command:
Pull the scale set name from the ID (it’ll look something likeaz aks nodepool show --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME --name SYSTEM_NODE_POOL_NAME --query 'scaleSetId' -o tsvaks-nodepool1-123456-vmss). - Stop and deallocate all VMs in the scale set
Run these commands to shut down the VMs—this stops charging you for their compute resources:az vmss stop --resource-group YOUR_RG_NAME --name SCALE_SET_NAME --no-wait az vmss deallocate --resource-group YOUR_RG_NAME --name SCALE_SET_NAME --no-wait - Restore your cluster on Monday
Start the VMs back up:
If the node doesn’t rejoin the cluster automatically, run this to refresh it:az vmss start --resource-group YOUR_RG_NAME --name SCALE_SET_NAME --no-waitaz aks nodepool upgrade --cluster-name YOUR_CLUSTER_NAME --resource-group YOUR_RG_NAME --name SYSTEM_NODE_POOL_NAME --node-count 1
Bonus: Automate this to skip manual work
To avoid doing this every weekend, use Azure Automation Runbooks to schedule the scale-down/stop commands for Friday evening, and scale-up/start commands for Monday morning. Here’s a quick breakdown:
- Create an Azure Automation account with a run-as account that has permissions to manage your AKS cluster and VM scale sets.
- Write two runbooks (PowerShell or Bash): one to scale user pools to 0 and stop system pool VMs, another to reverse the process.
- Schedule the runbooks to run at your desired off-hours.
Key things to keep in mind
- When scaling user pools to 0, make sure your workloads can handle being shut down over the weekend (use persistent storage if you need to retain data).
- Stopping system pool VMs will take your cluster offline—only do this if you don’t need access during off-hours.
- Always test these steps in a non-production cluster first to avoid unexpected issues.
内容的提问来源于stack exchange,提问作者Dario Romero

