You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AKS集群中kube-proxy修改后自动回滚的问题求助

问题描述

我正在沙箱AKS集群中进行实验调整,目标是将其配置为生产就绪状态。参考某书籍中AWS EKS环境下重新部署kube-proxy DaemonSet并修改配置的操作,我对AKS中的kube-proxy做了两处小修改:将CPU请求从100m调整为120m,将日志级别参数-v从3改为2。

问题是,修改后的DaemonSet和Pod在2-3分钟后会自动恢复到初始状态,执行回滚历史命令可见AKS在自动回滚:

> kubectl rollout history daemonset kube-proxy -n kube-system
daemonset.apps/kube-proxy 
REVISION  CHANGE-CAUSE
2         <none>
8         <none>
10        <none>
14        <none>
16        <none>

我通过声明式方式应用以下清单来重新部署:

---
apiVersion: apps/v1
kind: DaemonSet
metadata:
  labels:
    addonmanager.kubernetes.io/mode: Reconcile
    component: kube-proxy
    tier: node
    deployment: custom
  name: kube-proxy
  namespace: kube-system
spec:
  revisionHistoryLimit: 10
  selector:
    matchLabels:
      component: kube-proxy
      tier: node
  template:
    metadata:
      creationTimestamp: null
      labels:
        component: kube-proxy
        tier: node
        deployedBy: Luka
    spec:
      affinity:
        nodeAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            nodeSelectorTerms:
            - matchExpressions:
              - key: kubernetes.azure.com/cluster
                operator: Exists
              - key: type
                operator: NotIn
                values:
                - virtual-kubelet
              - key: kubernetes.io/os
                operator: In
                values:
                - linux
      containers:
      - command:
        - kube-proxy
        - --conntrack-max-per-core=0
        - --metrics-bind-address=0.0.0.0:10249
        - --kubeconfig=/var/lib/kubelet/kubeconfig
        - --cluster-cidr=10.244.0.0/16
        - --detect-local-mode=ClusterCIDR
        - --pod-interface-name-prefix=
        - --v=2
        image: mcr.microsoft.com/oss/kubernetes/kube-proxy:v1.23.12-hotfix.20220922.1
        imagePullPolicy: IfNotPresent
        name: kube-proxy
        resources:
          requests:
            cpu: 120m
        securityContext:
          privileged: true
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /var/lib/kubelet
          name: kubeconfig
          readOnly: true
        - mountPath: /etc/kubernetes/certs
          name: certificates
          readOnly: true
        - mountPath: /run/xtables.lock
          name: iptableslock
        - mountPath: /lib/modules
          name: modules
      dnsPolicy: ClusterFirst
      hostNetwork: true
      initContainers:
      - command:
        - /bin/sh
        - -c
        - |
          SYSCTL=/proc/sys/net/netfilter/nf_conntrack_max
          echo "Current net.netfilter.nf_conntrack_max: $(cat $SYSCTL)"
          DESIRED=$(awk -F= '/net.netfilter.nf_conntrack_max/ {print $2}' /etc/sysctl.d/999-sysctl-aks.conf)
          if [ -z "$DESIRED" ]; then
            DESIRED=$((32768*$(nproc)))
            if [ $DESIRED -lt 131072 ]; then
              DESIRED=131072
            fi

            echo "AKS custom config for net.netfilter.nf_conntrack_max not set."
            echo "Setting nf_conntrack_max to $DESIRED (32768 * $(nproc) cores, minimum 131072)."
            echo $DESIRED > $SYSCTL
          else
            echo "AKS custom config for net.netfilter.nf_conntrack_max set to $DESIRED."
            echo "Setting nf_conntrack_max to $DESIRED."
            echo $DESIRED > $SYSCTL
          fi
        image: mcr.microsoft.com/oss/kubernetes/kube-proxy:v1.23.12-hotfix.20220922.1
        imagePullPolicy: IfNotPresent
        name: kube-proxy-bootstrap
        resources:
          requests:
            cpu: 100m
        securityContext:
          privileged: true
        terminationMessagePath: /dev/termination-log
        terminationMessagePolicy: File
        volumeMounts:
        - mountPath: /etc/sysctl.d
          name: sysctls
        - mountPath: /lib/modules
          name: modules
      priorityClassName: system-node-critical
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext: {}
      terminationGracePeriodSeconds: 30
      tolerations:
      - key: CriticalAddonsOnly
        operator: Exists
      - effect: NoExecute
        operator: Exists
      - effect: NoSchedule
        operator: Exists
      volumes:
      - hostPath:
          path: /var/lib/kubelet
          type: ""
        name: kubeconfig
      - hostPath:
          path: /etc/kubernetes/certs
          type: ""
        name: certificates
      - hostPath:
          path: /run/xtables.lock
          type: FileOrCreate
        name: iptableslock
      - hostPath:
          path: /etc/sysctl.d
          type: Directory
        name: sysctls
      - hostPath:
          path: /lib/modules
          type: Directory
        name: modules
  updateStrategy:
    rollingUpdate:
      maxSurge: 0
      maxUnavailable: 1
    type: RollingUpdate
status:
  currentNumberScheduled: 4
  desiredNumberScheduled: 4
  numberAvailable: 4
  numberMisscheduled: 0
  numberReady: 4
  observedGeneration: 1
  updatedNumberScheduled: 4

我还尝试移除initContainer,以及参考帖子中的kubectl编辑DaemonSet方法,但均无效。请问我遗漏了什么?为何kube-proxy DaemonSet总是自动回滚?


原因分析与解决方案

核心原因

AKS集群中,kube-proxy是官方托管的核心组件,由AKS内置的addon-manager负责维护。你的DaemonSet元数据中保留了addonmanager.kubernetes.io/mode: Reconcile标签——这个标签会让addon-manager定期检查kube-proxy的配置,并强制同步回AKS的默认模板,直接覆盖你做的所有自定义修改,导致自动回滚。

解决办法

方法1:修改Addon管理模式(临时生效,集群升级后需重新配置)

将kube-proxy的addonmanager.kubernetes.io/mode标签值从Reconcile改为EnsureExists,这样addon-manager只会确保kube-proxy DaemonSet存在,不会干预你做的自定义配置:

  1. 编辑DaemonSet:
kubectl edit daemonset kube-proxy -n kube-system
  1. 在metadata.labels中找到addonmanager.kubernetes.io/mode,将值改为EnsureExists
  2. 保存退出后,重新应用你的自定义配置,此时修改不会再被自动回滚

方法2:使用AKS官方自定义kube-proxy配置(推荐,持久生效)

AKS支持通过官方CLI命令自定义kube-proxy配置,这是集群升级时也能保留的合规方式:

  1. 创建配置文件kube-proxy-config.yaml,包含你需要的参数:
apiVersion: kubeproxy.config.k8s.io/v1alpha1
kind: KubeProxyConfiguration
metricsBindAddress: 0.0.0.0:10249
clusterCIDR: 10.244.0.0/16
detectLocalMode: ClusterCIDR
verbosity: 2 # 修改日志级别为2
conntrack:
  maxPerCore: 0
  1. 通过Azure CLI将配置应用到集群:
az aks update --resource-group <你的资源组名称> --name <你的AKS集群名称> --kube-proxy-config kube-proxy-config.yaml
  1. 针对CPU请求的修改,完成上述配置后,再将addonmanager.kubernetes.io/mode改为EnsureExists,然后编辑DaemonSet调整resources.requests.cpu为120m,即可持久保留配置。

注意事项

  • 直接修改托管组件的DaemonSet属于非官方操作,AKS集群升级时可能会被重置,优先使用官方自定义配置方式。
  • 若你的AKS版本较旧,可能不支持--kube-proxy-config参数,此时只能通过修改Addon管理模式实现自定义配置。

内容的提问来源于stack exchange,提问作者Luka Klarić

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 17:15:14