You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Akka集群的K8s应用使用HPA时异常扩容且无法缩容求助

Akka集群应用K8s部署后HPA异常扩容至maxReplicas且无法缩容的问题解决

我们基于Akka集群开发的应用部署在Kubernetes集群中,使用HorizontalPodAutoscaler(HPA)实现自动扩缩容,但出现了刚部署完成、无负载的情况下直接扩容到maxReplicas,且始终无法缩容的问题。对应的StatefulSet和HPA配置如下:

apiVersion: apps/v1
kind: StatefulSet
metadata:
  name: app
  namespace: some-namespace
  labels:
    componentName: our-component
    app: our-component
    version: some-version
  annotations:
    prometheus.io/scrape: "true"
    prometheus.io/path: "/metrics"
    prometheus.io/port: "9252"
spec:
  serviceName: our-component
  replicas: 2
  selector:
    matchLabels:
      componentName: our-component
      app: our-app
  updateStrategy:
    type: RollingUpdate
  template:
    metadata:
      labels:
        componentName: our-component
        app: our-app
    spec:
      containers:
        - name: our-component-container
          image: image-path
          imagePullPolicy: Always
          resources:
            requests:
              cpu: .1
              memory: 500Mi
            limits:
              cpu: 1
              memory: 1Gi
          command:
            - "/microservice/bin/our-component"
          ports:
            - name: remoting
              containerPort: 8080
              protocol: TCP
          readinessProbe:
            httpGet:
              path: /ready
              port: 9085
            initialDelaySeconds: 40
            periodSeconds: 30
            failureThreshold: 3
            timeoutSeconds: 30
          livenessProbe:
            httpGet:
              path: /alive
              port: 9085
            initialDelaySeconds: 130
            periodSeconds: 30
            failureThreshold: 3
            timeoutSeconds: 5


apiVersion: autoscaling/v1
kind: HorizontalPodAutoscaler
metadata:
  name: app-hpa
  namespace: some-namespace
  labels:
    componentName: our-component
    app: our-app
spec:
  minReplicas: 2
  maxReplicas: 8
  scaleTargetRef:
    apiVersion: apps/v1
    kind: StatefulSet
    name: our-component
  targetCPUUtilizationPercentage: 75

核心问题排查与解决

1. HPA目标资源名称不匹配(最直接原因)

你的配置存在明显的名称不匹配问题:

  • StatefulSet的metadata.name为app
  • HPA的spec.scaleTargetRef.name为our-component

K8s的HPA需要精准匹配目标资源的名称,名称不匹配会导致HPA无法正确获取目标StatefulSet的Pod指标。当HPA无法获取有效CPU使用率数据时,会触发异常扩容行为,直接拉满到maxReplicas。

解决方案:修正HPA的scaleTargetRef.name为app,与StatefulSet名称保持一致:

spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: StatefulSet
    name: app  # 修改为与StatefulSet一致的名称

2. CPU指标采集异常

即使修正了名称,仍需确认指标采集是否正常:

  • 检查StatefulSet的Pod是否正常暴露/metrics端点,验证prometheus.io相关注解是否生效
  • 用kubectl top pods命令查看Pod的CPU实际使用率,确认K8s metrics-server能否正常采集到指标(HPA v1依赖该组件的CPU数据)
  • 确认Prometheus的采集配置(ServiceMonitor/PodMonitor)是否覆盖了目标namespace下的Pod

若指标采集异常,HPA同样会因无法获取有效数据出现异常扩缩容。

3. 扩缩容行为参数优化

如果是Akka集群初始化阶段的短暂负载波动导致误扩容,可以调整HPA的行为参数(建议使用HPA v2版本):

  • 配置扩缩容冷却时间,避免短时间内频繁调整副本数
  • 调整目标CPU使用率的容忍阈值,减少误触发

验证步骤

  1. 修正HPA配置后,执行kubectl apply -f <hpa配置文件路径>更新HPA
  2. 用kubectl describe hpa app-hpa查看HPA状态,确认Target和Current CPU使用率是否正常显示
  3. 等待默认缩容冷却时间(5分钟),观察Pod数量是否回落到minReplicas值

内容的提问来源于stack exchange,提问作者Karan Khanna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 04:54:34