You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免Spring应用启动时K8s HPA因CPU突增触发扩容

问题描述

在K8s 1.26.3上管理Spring应用,Pod启动运行需2~3分钟,探针等待时间设置为4分钟。重新部署应用后,Pod初始化阶段会出现CPU突增现象,导致滚动更新时触发HPA扩容。需求是在应用初始化阶段不收集CPU指标,避免误扩容。

K8s版本:1.26.3

Deployment核心配置

...
livenessProbe:
  failureThreshold: 3
  httpGet:
    path: /probex
    port: 9006
    scheme: HTTP
  initialDelaySeconds: 240
  periodSeconds: 60
  successThreshold: 1
  timeoutSeconds: 1
name: sample-container
ports:
- containerPort: 9006
  name: http
  protocol: TCP
readinessProbe:
  failureThreshold: 3
  httpGet:
    path: /probex
    port: 9006
    scheme: HTTP
  initialDelaySeconds: 240
  periodSeconds: 60
  successThreshold: 1
  timeoutSeconds: 1
resources:
  limits:
    memory: 8Gi
  requests:
    cpu: 500m
    memory: 8Gi
...

HPA核心配置

...
spec:
  maxReplicas: 5
  metrics:
  - resource:
      name: memory
      target:
        averageUtilization: 70
        type: Utilization
    type: Resource
  - resource:
      name: cpu
      target:
        averageUtilization: 70
        type: Utilization
    type: Resource
  minReplicas: 1
...
可行解决方案

1. 调整HPA扩容稳定窗口,过滤短暂CPU突增

通过HPA的behavior字段配置扩容稳定窗口,让HPA忽略短暂的CPU峰值,只有当高负载持续一段时间后才触发扩容。适合不想改动指标收集链路的场景。

修改HPA配置,添加scaleUp的稳定窗口:

spec:
  ...
  behavior:
    scaleUp:
      stabilizationWindowSeconds: 300  # 5分钟,覆盖Pod初始化的CPU突增周期
      policies:
      - type: Percent
        value: 100
        periodSeconds: 60
  • stabilizationWindowSeconds设置为300秒(5分钟),确保HPA只会基于5分钟内的平均CPU利用率判断是否扩容,过滤初始化阶段的短暂突增。
  • 该配置不影响指标收集,但会延迟扩容决策,避免误触发。

2. 基于Pod就绪状态过滤CPU指标(需Prometheus Adapter)

如果使用Prometheus作为指标数据源,可通过Prometheus Adapter自定义CPU指标,只收集**就绪状态(Ready=True)**的Pod的CPU利用率,彻底排除初始化阶段未就绪Pod的指标。

步骤:

  1. 编写Prometheus查询,过滤就绪Pod的CPU使用率:

    sum(irate(container_cpu_usage_seconds_total{namespace="你的命名空间", pod=~"sample-.*", container="sample-container"}[1m])) by (pod)
    /
    sum(kube_pod_container_resource_requests{namespace="你的命名空间", pod=~"sample-.*", container="sample-container", resource="cpu"}) by (pod)
    * 100
    

    结合kube_pod_status_ready{condition="true"}标签,确保只计算就绪Pod的指标。

  2. 在Prometheus Adapter的配置中添加该自定义指标,映射为HPA可用的resource类型指标。

  3. 修改HPA配置,使用自定义的CPU指标替代原生resource cpu指标。

3. 调整Pod就绪探针逻辑,提前标记就绪状态(可选)

如果应用在2~3分钟后已经具备服务能力,可将readinessProbe的initialDelaySeconds调整为180秒(3分钟),让Pod提前进入就绪状态,HPA只会计算就绪Pod的指标(autoscaling/v2默认行为)。但需确保此时应用确实能正常处理请求,避免影响服务可用性。

修改Deployment的readiness探针:

readinessProbe:
  ...
  initialDelaySeconds: 180  # 从240秒改为180秒
  • 该方案依赖应用实际就绪时间,需测试确认应用在3分钟后能稳定响应探针请求。

内容的提问来源于stack exchange,提问作者Daniel Lee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 11:17:43