You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GKE中基于资源动态调整指定Deployment副本数的方案

基于GKE集群资源动态调整指定Deployment副本数的方案

核心需求梳理

  • 针对GKE集群中部分跨命名空间的Deployment,根据集群可用资源(重点是内存)动态调整副本数
  • 正常状态下保持默认副本数(示例为3),集群资源不足时自动缩容,部分Deployment可缩至0
  • 仅对标记的特定Deployment生效,不影响其他重要服务

实现方案

1. 标记目标Deployment

先给需要动态调整的Deployment添加统一标签,方便后续识别,比如scalable: "true"标记可缩容,priority: "low"标记低优先级,scale-to-zero: "true"标记允许缩至0。修改后的Deployment YAML示例:

apiVersion: apps/v1
kind: Deployment
metadata:
  labels:
    app: my-deployment
    scalable: "true"       # 标记为可动态缩容的部署
    priority: "low"        # 低优先级,资源不足时优先缩容
    scale-to-zero: "true"  # 允许缩容至0(按需添加)
  name: my-deployment
  namespace: my-namespace
spec:
  replicas: 3 # 默认副本数
  selector:
    matchLabels:
      app: my-deployment
  template:
    metadata:
      labels:
        app: my-deployment
        scalable: "true"
        priority: "low"
        scale-to-zero: "true"
    spec:
      containers:
      - image: my/image:version
        name: my-deployment
        resources:
          requests:        # 必须定义资源请求,方便集群计算可用资源
            memory: "256Mi"
            cpu: "100m"
          limits:
            memory: "512Mi"
            cpu: "200m"

2. 自定义脚本+CronJob实现动态缩容

通过定期检查集群资源的脚本,结合Kubernetes CronJob实现自动化调整,这是最灵活的方案:

步骤1:编写资源检查与缩容脚本

创建Shell脚本,核心逻辑是判断集群可用内存阈值,然后批量调整目标Deployment的副本数:

#!/bin/bash

# 获取集群总内存与已用内存(单位:Mi)
TOTAL_MEM=$(kubectl top nodes --no-headers | awk '{sum+=$4} END {print sum}')
USED_MEM=$(kubectl top nodes --no-headers | awk '{sum+=$3} END {print sum}')
AVAILABLE_MEM=$((TOTAL_MEM - USED_MEM))
# 设定资源阈值:可用内存低于总内存20%时触发缩容
THRESHOLD_MEM=$((TOTAL_MEM * 20 / 100))

if [ $AVAILABLE_MEM -lt $THRESHOLD_MEM ]; then
  echo "集群内存不足,开始缩容低优先级Deployment..."
  # 缩容允许到0的Deployment
  kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero=true -o name | while read deploy; do
    kubectl scale $deploy --replicas=0
  done
  # 缩容其他低优先级Deployment到1个副本
  kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero!=true -o name | while read deploy; do
    kubectl scale $deploy --replicas=1
  done
else
  echo "集群资源充足,恢复Deployment默认副本数..."
  # 恢复允许到0的Deployment到默认3个副本
  kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero=true -o name | while read deploy; do
    kubectl scale $deploy --replicas=3
  done
  # 恢复其他低优先级Deployment到默认3个副本
  kubectl get deployments --all-namespaces -l scalable=true,scale-to-zero!=true -o name | while read deploy; do
    kubectl scale $deploy --replicas=3
  done
fi

步骤2:用CronJob定期执行脚本

使用包含kubectl的镜像,创建CronJob定期运行脚本(示例为每5分钟检查一次):

apiVersion: batch/v1
kind: CronJob
metadata:
  name: deployment-scaler
  namespace: kube-system
spec:
  schedule: "*/5 * * * *" # 每5分钟执行一次
  jobTemplate:
    spec:
      template:
        spec:
          serviceAccountName: deployment-scaler-sa # 需要赋予操作Deployment的权限
          containers:
          - name: scaler
            image: bitnami/kubectl:latest
            command: ["bash", "-c"]
            args:
            - |
              # 粘贴上面的脚本内容
          restartPolicy: OnFailure

步骤3:配置ServiceAccount权限

创建具备跨命名空间操作Deployment和读取节点资源权限的ServiceAccount:

apiVersion: v1
kind: ServiceAccount
metadata:
  name: deployment-scaler-sa
  namespace: kube-system

---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: deployment-scaler-role
rules:
- apiGroups: ["apps"]
  resources: ["deployments"]
  verbs: ["get", "list", "update", "patch"]
- apiGroups: [""]
  resources: ["nodes"]
  verbs: ["get", "list"]

---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: deployment-scaler-binding
subjects:
- kind: ServiceAccount
  name: deployment-scaler-sa
  namespace: kube-system
roleRef:
  kind: ClusterRole
  name: deployment-scaler-role
  apiGroup: rbac.authorization.k8s.io

3. 替代方案说明

GKE自带的Cluster Autoscaler和Vertical Pod Autoscaler主要用于节点扩缩容和Pod资源调整,无法实现按优先级缩容Deployment副本数,也不支持将副本缩至0,因此更适合用自定义脚本方案。

注意事项

  • 必须给Deployment定义resources.requests和resources.limits,否则无法准确计算集群资源使用情况
  • 可根据实际需求调整脚本中的资源阈值、副本数恢复策略
  • 测试阶段可降低CronJob执行频率,避免频繁调整影响服务

内容的提问来源于stack exchange,提问作者dan1st

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 00:21:26