AKS自定义调度器未按预期实现Bin packing装箱功能
自定义K8s调度器(AKS)未按MostAllocated策略装箱问题排查
背景
在AKS上部署了自定义调度器,通过OTA Gatekeeper自动将所有Pod的调度器切换为该自定义调度器,目标是实现高效Bin packing装箱,配合de-scheduler移除低利用率节点,降低集群整体成本。调度器与Gatekeeper部署正常,Pod已使用自定义调度器调度,但测试结果不符合预期。
测试场景
- AKS集群包含3个几乎为空的节点
- 存在名为
clustercheck的Deployment,其Pod请求10m CPU - 将Deployment副本数调整为10后,副本均匀分布在3个节点上;预期基于MostAllocated算法,所有副本应被调度到同一节点
调度器配置
- schedulerName: agys-scheduler pluginConfig: - args: apiVersion: kubescheduler.config.k8s.io/v1beta2 kind: NodeResourcesFitArgs scoringStrategy: resources: - name: cpu weight: 1 - name: memory weight: 1 type: MostAllocated name: NodeResourcesFit plugins: score: enabled: - name: NodeResourcesFit weight: 1
实际Pod分布
clustercheck-567c5c49f6-8g7c9 1/1 Running 0 26s 10.240.214.54 aks-subnet02-10595059-vmss000006 <none> <none> clustercheck-567c5c49f6-97b9p 1/1 Running 0 26s 10.240.214.9 aks-subnet02-10595059-vmss000006 <none> <none> clustercheck-567c5c49f6-n8n8v 1/1 Running 0 26s 10.240.214.90 aks-subnet02-10595059-vmss000009 <none> <none> clustercheck-567c5c49f6-njzdt 1/1 Running 0 26s 10.240.214.65 aks-subnet02-10595059-vmss000009 <none> <none> clustercheck-567c5c49f6-p6cfw 1/1 Running 0 5d18h 10.240.214.7 aks-subnet02-10595059-vmss000006 <none> <none> clustercheck-567c5c49f6-rh9zz 1/1 Running 0 26s 10.240.214.42 aks-subnet02-10595059-vmss000006 <none> <none> clustercheck-567c5c49f6-rqtrm 1/1 Running 0 26s 10.240.214.73 aks-subnet02-10595059-vmss000009 <none> <none> clustercheck-567c5c49f6-s84x4 1/1 Running 0 26s 10.240.214.5 aks-subnet02-10595059-vmss000006 <none> <none> clustercheck-567c5c49f6-vzcn9 1/1 Running 0 26s 10.240.214.110 aks-subnet02-10595059-vmss000007 <none> <none> clustercheck-567c5c49f6-x7ssx 1/1 Running 0 26s 10.240.214.8 aks-subnet02-10595059-vmss000006 <none> <none> webgoat-66d8d7cb57-7pl46 1/1 Running 0 5d18h 10.240.214.57 aks-subnet02-10595059-vmss000006 <none> <none>
调度器日志信息
I1115 02:54:46.222071 1 framework.go:478] "MultiPoint plugin is explicitly re-configured; overriding" plugin="NodeResourcesFit" I1115 02:54:46.224457 1 configfile.go:102] "Using component config" config=< apiVersion: kubescheduler.config.k8s.io/v1 clientConnection: acceptContentTypes: "" burst: 100 contentType: application/vnd.kubernetes.protobuf kubeconfig: "" qps: 50 enableContentionProfiling: true enableProfiling: true kind: KubeSchedulerConfiguration leaderElection: leaderElect: false leaseDuration: 15s renewDeadline: 10s resourceLock: leases resourceName: kube-scheduler resourceNamespace: kube-system retryPeriod: 2s parallelism: 16 percentageOfNodesToScore: 0 podInitialBackoffSeconds: 1 podMaxBackoffSeconds: 10 profiles: - pluginConfig: - args: apiVersion: kubescheduler.config.k8s.io/v1 kind: DefaultPreemptionArgs minCandidateNodesAbsolute: 100 minCandidateNodesPercentage: 10 name: DefaultPreemption - args: apiVersion: kubescheduler.config.k8s.io/v1 hardPodAffinityWeight: 1 kind: InterPodAffinityArgs name: InterPodAffinity - args: apiVersion: kubescheduler.config.k8s.io/v1 kind: NodeAffinityArgs name: NodeAffinity - args: apiVersion: kubescheduler.config.k8s.io/v1 kind: NodeResourcesBalancedAllocationArgs resources: - name: cpu weight: 1 - name: memory weight: 1 name: NodeResourcesBalancedAllocation - args: apiVersion: kubescheduler.config.k8s.io/v1 kind: NodeResourcesFitArgs scoringStrategy: resources: - name: cpu weight: 100 - name: memory weight: 100 type: MostAllocated name: NodeResourcesFit - args: apiVersion: kubescheduler.config.k8s.io/v1 defaultingType: System kind: PodTopologySpreadArgs name: PodTopologySpread - args: apiVersion: kubescheduler.config.k8s.io/v1 bindTimeoutSeconds: 600 kind: VolumeBindingArgs name: VolumeBinding plugins: bind: {} filter: {} multiPoint: enabled: - name: PrioritySort weight: 0 - name: NodeUnschedulable weight: 0 - name: NodeName weight: 0 - name: TaintToleration weight: 3 - name: NodeAffinity weight: 2 - name: NodePorts weight: 0 - name: NodeResourcesFit weight: 1 - name: VolumeRestrictions weight: 0 - name: EBSLimits weight: 0 - name: GCEPDLimits weight: 0 - name: NodeVolumeLimits weight: 0 - name: AzureDiskLimits weight: 0 - name: VolumeBinding weight: 0 - name: VolumeZone weight: 0 - name: PodTopologySpread weight: 2 - name: InterPodAffinity weight: 2 - name: DefaultPreemption weight: 0 - name: NodeResourcesBalancedAllocation weight: 1 - name: ImageLocality weight: 1 - name: DefaultBinder weight: 0 permit: {} postBind: {} postFilter: {} preBind: {} preFilter: {} preScore: {} queueSort: {} reserve: {} score: enabled: - name: NodeResourcesFit weight: 100 schedulerName: agys-scheduler
问题原因与修复方案
1. 冲突插件干扰
- PodTopologySpread插件:日志显示该插件在multiPoint阶段权重为2,其核心逻辑是将Pod均匀分布到不同拓扑域(默认按节点划分),直接与MostAllocated的装箱目标冲突,导致Pod被分散调度。
- NodeResourcesBalancedAllocation插件:该插件同样在multiPoint阶段启用,权重为1,会倾向于平衡节点资源利用率,与装箱策略相悖。
2. 配置优先级问题
日志提示NodeResourcesFit被MultiPoint配置覆盖,且原配置中NodeResourcesFit的权重过低(仅1),无法抵消其他插件的影响。
3. 修改后的调度器配置
- schedulerName: agys-scheduler pluginConfig: - args: apiVersion: kubescheduler.config.k8s.io/v1 kind: NodeResourcesFitArgs scoringStrategy: resources: - name: cpu weight: 100 - name: memory weight: 100 type: MostAllocated name: NodeResourcesFit plugins: multiPoint: enabled: - name: PrioritySort weight: 0 - name: NodeUnschedulable weight: 0 - name: NodeName weight: 0 - name: TaintToleration weight: 3 - name: NodeAffinity weight: 2 - name: NodePorts weight: 0 - name: NodeResourcesFit weight: 100 - name: VolumeRestrictions weight: 0 - name: AzureDiskLimits weight: 0 - name: VolumeBinding weight: 0 - name: InterPodAffinity weight: 2 - name: DefaultPreemption weight: 0 - name: ImageLocality weight: 1 - name: DefaultBinder weight: 0 disabled: - name: PodTopologySpread - name: NodeResourcesBalancedAllocation score: enabled: - name: NodeResourcesFit weight: 100
4. 验证步骤
- 应用修改后的调度器配置并重启自定义调度器
- 删除现有
clustercheckPod,重新触发调度 - 检查Pod分布是否集中到同一节点
内容的提问来源于stack exchange,提问作者Hari Narayanan
相关产品推荐
相关产品推荐

