GKE中基于资源动态调整指定Deployment副本数的方案
基于GKE集群资源动态调整指定Deployment副本数的方案
核心需求梳理
- 针对GKE集群中部分跨命名空间的Deployment,根据集群可用资源(重点是内存)动态调整副本数
- 正常状态下保持默认副本数(示例为3),集群资源不足时自动缩容,部分Deployment可缩至0
- 仅对标记的特定Deployment生效,不影响其他重要服务
实现方案
1. 标记目标Deployment
先给需要动态调整的Deployment添加统一标签,方便后续识别,比如scalable: "true"标记可缩容,priority: "low"标记低优先级,scale-to-zero: "true"标记允许缩至0。修改后的Deployment YAML示例:
apiVersion: apps/v1 kind: Deployment metadata: labels: app: my-deployment scalable: "true" # 标记为可动态缩容的部署 priority: "low" # 低优先级,资源不足时优先缩容 scale-to-zero: "true" # 允许缩容至0(按需添加) name: my-deployment namespace: my-namespace spec: replicas: 3 # 默认副本数 selector: matchLabels: app: my-deployment template: metadata: labels: app: my-deployment scalable: "true" priority: "low" scale-to-zero: "true" spec: containers: - image: my/image:version name: my-deployment resources: requests: # 必须定义资源请求,方便集群计算可用资源 memory: "256Mi" cpu: "100m" limits: memory: "512Mi" cpu: "200m"
2. 自定义脚本+CronJob实现动态缩容
通过定期检查集群资源的脚本,结合Kubernetes CronJob实现自动化调整,这是最灵活的方案:
步骤1:编写资源检查与缩容脚本
创建Shell脚本,核心逻辑是判断集群可用内存阈值,然后批量调整目标Deployment的副本数:
#!/bin/bash # 获取集群总内存与已用内存(单位:Mi) TOTAL_MEM=$(kubectl top nodes --no-headers | awk '{sum+=$4} END {print sum}') USED_MEM=$(kubectl top nodes --no-headers | awk '{sum+=$3} END {print sum}') AVAILABLE_MEM=$((TOTAL_MEM - USED_MEM)) # 设定资源阈值:可用内存低于总内存20%时触发缩容 THRESHOLD_MEM=$((TOTAL_MEM * 20 / 100)) if [ $AVAILABLE_MEM -lt $THRESHOLD_MEM ]; then echo "集群内存不足,开始缩容低优先级Deployment..." # 缩容允许到0的Deployment kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero=true -o name | while read deploy; do kubectl scale $deploy --replicas=0 done # 缩容其他低优先级Deployment到1个副本 kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero!=true -o name | while read deploy; do kubectl scale $deploy --replicas=1 done else echo "集群资源充足,恢复Deployment默认副本数..." # 恢复允许到0的Deployment到默认3个副本 kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero=true -o name | while read deploy; do kubectl scale $deploy --replicas=3 done # 恢复其他低优先级Deployment到默认3个副本 kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero!=true -o name | while read deploy; do kubectl scale $deploy --replicas=3 done fi
步骤2:用CronJob定期执行脚本
使用包含kubectl的镜像,创建CronJob定期运行脚本(示例为每5分钟检查一次):
apiVersion: batch/v1 kind: CronJob metadata: name: deployment-scaler namespace: kube-system spec: schedule: "*/5 * * * *" # 每5分钟执行一次 jobTemplate: spec: template: spec: serviceAccountName: deployment-scaler-sa # 需要赋予操作Deployment的权限 containers: - name: scaler image: bitnami/kubectl:latest command: ["bash", "-c"] args: - | # 粘贴上面的脚本内容 restartPolicy: OnFailure
步骤3:配置ServiceAccount权限
创建具备跨命名空间操作Deployment和读取节点资源权限的ServiceAccount:
apiVersion: v1 kind: ServiceAccount metadata: name: deployment-scaler-sa namespace: kube-system --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRole metadata: name: deployment-scaler-role rules: - apiGroups: ["apps"] resources: ["deployments"] verbs: ["get", "list", "update", "patch"] - apiGroups: [""] resources: ["nodes"] verbs: ["get", "list"] --- apiVersion: rbac.authorization.k8s.io/v1 kind: ClusterRoleBinding metadata: name: deployment-scaler-binding subjects: - kind: ServiceAccount name: deployment-scaler-sa namespace: kube-system roleRef: kind: ClusterRole name: deployment-scaler-role apiGroup: rbac.authorization.k8s.io
3. 替代方案说明
GKE自带的Cluster Autoscaler和Vertical Pod Autoscaler主要用于节点扩缩容和Pod资源调整,无法实现按优先级缩容Deployment副本数,也不支持将副本缩至0,因此更适合用自定义脚本方案。
注意事项
- 必须给Deployment定义
resources.requests和resources.limits,否则无法准确计算集群资源使用情况 - 可根据实际需求调整脚本中的资源阈值、副本数恢复策略
- 测试阶段可降低CronJob执行频率,避免频繁调整影响服务
内容的提问来源于stack exchange,提问作者dan1st
相关产品推荐
相关产品推荐

