You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为AKS中使用LRS磁盘的StatefulSet配置区域亲和调度?

解决AKS StatefulSet Pod与LRS磁盘跨区调度Pending问题

1. 修正StorageClass的核心配置错误

你的StorageClass存在两处关键拼写错误,导致拓扑感知绑定无法正常工作:

  • 将provisioned字段修正为provisioner(Kubernetes StorageClass标准字段,用于指定CSI驱动)
  • 将WaitForFirstCustomer修正为WaitForFirstConsumer(标准延迟绑定模式,确保PV在Pod调度完成后创建,并绑定到Pod所在的可用区)

修正后的完整StorageClass配置:

allowVolumeExpansion: true
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  annotations:
    storageclass.kubernetes.io/is-default-class: true
  name: my-block-sc
parameters:
  cachingmode: ReadOnly
  diskEncryptionSetID: ...
  diskEncryptionType: EncryptionAtRestWithCustomerKey
  networkAccessPolicy: DenyAll
  skuName: StandardSSD_LRS
provisioner: disk.csi.azure.com
reclaimPolicy: Retain 
volumeBindingMode: WaitForFirstConsumer

2. 配置StatefulSet的拓扑感知调度

当StorageClass配置正确后,结合StatefulSet的特性,即可实现Pod重启后自动调度到PV所在的可用区:

  • StatefulSet的每个Pod会关联唯一命名的PVC(格式为<statefulset-name>-<pod-index>),该PVC会永久绑定到固定PV
  • Azure Disk CSI驱动会自动为PV添加拓扑标签topology.disk.csi.azure.com/zone,对应磁盘所在的可用区
  • Kubernetes调度器会识别PV的拓扑约束,在Pod重启时自动将其调度到同一可用区的节点上

如果需要显式强化调度规则,可以在StatefulSet的Pod模板中添加节点亲和配置,匹配节点的可用区标签(与PV拓扑标签一致):

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: your-statefulset-name
spec:
  serviceName: your-service-name
  replicas: 3
  selector:
    matchLabels:
      app: your-app-label
  template:
    metadata:
      labels:
        app: your-app-label
    spec:
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: topology.kubernetes.io/zone
                operator: In
                values:
                # 可手动填写集群支持的可用区,比如:
                # - switzerlandnorth-1
                # - switzerlandnorth-2
                # - switzerlandnorth-3
      volumeClaimTemplates:
      - metadata:
          name: data
        spec:
          accessModes: ["ReadWriteOnce"]
          storageClassName: my-block-sc
          resources:
            requests:
              storage: 10Gi

3. 验证配置效果

部署修正后的资源后,可通过以下步骤验证:

  1. 查看PV的拓扑标签:kubectl describe pv <pv-name>,确认topology.disk.csi.azure.com/zone标签存在
  2. 查看Pod调度节点:kubectl get pods -o wide,确认Pod运行在PV所在可用区的节点上
  3. 手动删除Pod:kubectl delete pod <pod-name>,观察重启后的Pod是否调度到同一可用区的节点

此方案既保留了跨可用区的高可用性(StatefulSet Pod可分布在不同可用区),又避免了Pod因跨区调度导致的Pending状态。

内容的提问来源于stack exchange,提问作者Matthias Güntert

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 14:52:49