You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于KEDA实现Pod高效缩容:缩容前需检查用户活动

实现K8s缩容时等待Pod无活跃用户的方案结合preStop与readinessProbe

针对你的Spring Boot + KEDA + RabbitMQ场景,要实现缩容前等待Pod无用户活跃请求,核心是通过readinessProbe切断新请求流入 + preStop钩子等待现有请求处理完成,以下是具体落地步骤:


一、Spring Boot应用内添加状态控制逻辑

首先需要在应用中实现两个核心能力:统计活跃用户请求数、响应终止信号并等待请求结束。

1. 实现活跃请求统计与状态接口

通过Filter拦截用户请求,配合原子变量统计活跃请求数,同时提供就绪状态和预终止接口:

import jakarta.servlet.*;
import jakarta.servlet.annotation.WebFilter;
import org.springframework.beans.factory.annotation.Autowired;
import org.springframework.web.bind.annotation.GetMapping;
import org.springframework.web.bind.annotation.PostMapping;
import org.springframework.web.bind.annotation.RestController;

import java.io.IOException;
import java.util.concurrent.atomic.AtomicBoolean;
import java.util.concurrent.atomic.AtomicInteger;

@RestController
public class InstanceStatusController {
    private final AtomicInteger activeRequests = new AtomicInteger(0);
    private final AtomicBoolean isTerminating = new AtomicBoolean(false);

    // 暴露活跃请求数(用于调试,非必须)
    @GetMapping("/actives")
    public int getActiveRequests() {
        return activeRequests.get();
    }

    // 给K8s readinessProbe用的就绪状态接口
    @GetMapping("/ready")
    public boolean isReady() {
        // 处于终止流程时,返回false,让K8s停止转发新请求
        return !isTerminating.get();
    }

    // 接收preStop钩子的终止信号
    @PostMapping("/pre-stop")
    public void preStop() throws InterruptedException {
        isTerminating.set(true);
        // 循环等待活跃请求数降为0,最多等待terminationGracePeriodSeconds时长
        while (activeRequests.get() > 0) {
            Thread.sleep(1000);
        }
    }

    // 活跃请求计数+1(Filter调用)
    public void incrementActive() {
        activeRequests.incrementAndGet();
    }

    // 活跃请求计数-1(Filter调用)
    public void decrementActive() {
        activeRequests.decrementAndGet();
    }
}

// 拦截所有用户请求路径(替换为你的业务接口前缀)
@WebFilter(urlPatterns = "/api/*")
class ActiveRequestFilter implements Filter {
    @Autowired
    private InstanceStatusController statusController;

    @Override
    public void doFilter(ServletRequest request, ServletResponse response, FilterChain chain) throws IOException, ServletException {
        try {
            statusController.incrementActive();
            chain.doFilter(request, response);
        } finally {
            statusController.decrementActive();
        }
    }
}

二、配置Kubernetes Deployment

在Deployment的Pod模板中,配置readinessProbe和preStop钩子,对接上面的接口:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: your-spring-boot-app
spec:
  template:
    spec:
      containers:
      - name: app-container
        image: your-app-image:tag
        ports:
        - containerPort: 8080
        # 就绪探针:判断Pod是否可以接收新请求
        readinessProbe:
          httpGet:
            path: /ready
            port: 8080
          initialDelaySeconds: 5
          periodSeconds: 5
          failureThreshold: 1
        # 预终止钩子:触发应用进入终止等待流程
        lifecycle:
          preStop:
            httpGet:
              path: /pre-stop
              port: 8080
        # 终止宽限期:设置足够长的时间,确保所有请求能处理完成(建议5-10分钟)
        terminationGracePeriodSeconds: 300

三、优化KEDA缩容策略

在ScaledObject中配置缩容冷却时间,避免KEDA频繁触发缩容,给Pod足够的时间处理剩余请求:

apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: rabbitmq-scaler
spec:
  scaleTargetRef:
    name: your-spring-boot-app
  minReplicaCount: 1
  maxReplicaCount: 10
  # 缩容冷却时间:比如5分钟,确保队列消息处理稳定后再缩容
  cooldownPeriod: 300
  triggers:
  - type: rabbitmq
    metadata:
      queueName: your-queue-name
      host: amqp://rabbitmq-service:5672
      queueLength: "X" # 你的扩容阈值

四、关键注意事项

  • 终止宽限期设置:terminationGracePeriodSeconds必须大于业务请求的最大处理时长,避免K8s强制杀死仍在处理请求的Pod。
  • 请求覆盖范围:确保Filter/AOP覆盖所有用户请求路径,包括异步请求、WebSocket连接等(如果有),避免遗漏活跃请求计数。
  • 预终止钩子超时处理:可以在preStop的等待逻辑中添加最大等待次数限制,防止因异常请求导致无限阻塞。
  • 测试验证:模拟用户持续请求时触发缩容,查看Pod是否会等待请求完成后再终止,同时验证新请求是否会被转发到其他就绪Pod。

内容的提问来源于stack exchange,提问作者Mohamed Nalouti

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 12:55:02