You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure AKS节点池无法扩容求助:KEDA已扩Pod但节点未自动缩放

Alright, let's dig into why your AKS node pool isn't auto-scaling even though KEDA is working perfectly for your Celery pods. I've run into this exact scenario a few times, so here's a step-by-step breakdown to troubleshoot:

1. First, Confirm Cluster Autoscaler is Enabled & Properly Configured

AKS node auto-scaling relies on the Cluster Autoscaler (CA) component—so let's start by verifying it's set up correctly:

  • Run this Azure CLI command to check your node pool's auto-scaling settings:
    az aks show --resource-group <your-resource-group> --name <your-aks-cluster> --query "agentPoolProfiles[].{name:name, enableAutoScaling:enableAutoScaling, minCount:minCount, maxCount:maxCount}"
    
    Double-check that enableAutoScaling is true, and maxCount is higher than your current number of nodes. If maxCount equals your current node count, CA can't scale up.
  • Verify the Cluster Autoscaler pod is running in your cluster:
    kubectl get pods -n kube-system -l app=cluster-autoscaler
    
    If the pod isn't in a Running state, check its logs for errors:
    kubectl logs -n kube-system <cluster-autoscaler-pod-name>
    
2. Check Pod Scheduling Constraints

Sometimes the issue isn't with CA itself—it's that your Celery pods can't be scheduled onto new nodes (or CA can't calculate the resource need). Here's what to check:

  • Resource Requests: The Cluster Autoscaler uses pod resource requests to determine if it needs to add nodes. If your Celery worker pods don't have resources.requests defined, CA might not recognize the resource shortage. Verify this with:
    kubectl describe pod <celery-worker-pod-name> -n <your-namespace> | grep -A5 Resources
    
    Always set reasonable CPU/memory requests for your pods—this is critical for CA to work.
  • Node Selectors/Taints & Tolerations: If your pods have a node selector, make sure your node pool has the matching labels. Similarly, if your nodes have taints, ensure your pods have corresponding tolerations. Check with:
    # Check pod's node selector/tolerations
    kubectl describe pod <celery-worker-pod-name> -n <your-namespace> | grep -A10 "Node Selector\|Tolerations"
    # Check node's labels/taints
    kubectl describe node <your-node-name> | grep -A10 "Labels\|Taints"
    
  • Resource Quotas/Pod Disruption Budgets: Ensure your namespace doesn't have a resource quota that's limiting pod count, and no PDB is blocking node scaling:
    kubectl get quota -n <your-namespace>
    kubectl get pdb -n <your-namespace>
    
3. Dig Into Cluster Autoscaler Logs for Clues

The CA logs will tell you exactly why it's not scaling up. Look for messages like:

  • No candidates for scale up: CA can't find a node pool that can accommodate the pending pods.
  • Insufficient capacity: Azure doesn't have enough VM resources in the region (this ties to subscription quotas).
  • Max nodes count reached: Your node pool's maxCount is already at the limit.
  • Pod has unmet node selector/taint: The pod can't be scheduled on any existing or potential new nodes.

To pull the logs (filter for scale-up events):

kubectl logs -n kube-system <cluster-autoscaler-pod-name> | grep -i "scale up\|no candidates\|insufficient"
4. Check Azure Subscription & VNet Limits

Sometimes the bottleneck is outside the cluster:

  • VM Quotas: Your Azure subscription might have a quota limit on the VM size used by your node pool. Check this with:
    az vm list-usage --location <your-cluster-region> --query "[?localName=='<your-vm-family> Family vCPUs']"
    
    (Replace <your-vm-family> with something like Standard DSv3.)
  • VNet Subnet IPs: If your node pool's subnet has no available IP addresses, CA can't provision new nodes. Check subnet usage:
    az network vnet subnet show --resource-group <your-resource-group> --vnet-name <your-vnet-name> --name <your-subnet-name> --query "{addressPrefix:addressPrefix, usedIPs:ipConfigurations.length}"
    
    Calculate remaining IPs (subtract usedIPs from total available in the prefix).
5. Test Manual Node Scaling

To isolate the issue, try manually scaling your node pool:

az aks scale --resource-group <your-resource-group> --name <your-aks-cluster> --node-count <higher-count> --nodepool-name <your-node-pool-name>
  • If manual scaling fails: The problem is with Azure resources (quota, VNet, permissions).
  • If manual scaling works: The issue is with the Cluster Autoscaler configuration or pod scheduling logic (go back to steps 2-3).

If you're still stuck, share the relevant Cluster Autoscaler logs and your node pool configuration details, and we can narrow it down further.

内容的提问来源于stack exchange,提问作者Paramesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:13:59