You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

启用Liveness与Readiness探针后K8S Pod进入CrashLoopBackOff的原因

问题分析与解决方案

核心原因

你的Pod进入CrashLoopBackOff的关键问题是存活探针(Liveness Probe)的默认配置过于激进:

  • Kubernetes中存活探针默认initialDelaySeconds=0,即Pod启动后立刻开始检查健康状态
  • Spring Boot应用启动需要加载依赖、初始化上下文等时间,此时/actuator/health/liveness端点尚未就绪,存活探针检测失败触发Pod重启
  • 反复重启后就进入了CrashLoopBackOff状态

而单独使用就绪探针时,就绪探针失败只会将Pod从Service的端点列表中移除,不会触发重启,因此部署能正常进行。

正确配置建议

1. 优化探针参数

给存活探针添加启动延迟,同时调整检测周期、超时和失败阈值,适配应用的实际启动速度:

livenessProbe:
  httpGet:
    port: 8000
    path: /actuator/health/liveness
  initialDelaySeconds: 30  # 应用启动后30秒再开始检测,根据实际启动时间调整
  periodSeconds: 10        # 每10秒检测一次
  timeoutSeconds: 5        # 检测超时时间5秒
  failureThreshold: 3      # 连续3次失败才判定为不健康
readinessProbe:
  httpGet:
    port: 8000
    path: /actuator/health/readiness
  initialDelaySeconds: 15  # 就绪探针可稍早开始检测
  periodSeconds: 5
  timeoutSeconds: 3
  failureThreshold: 2

提示:可以通过kubectl logs <pod-name>查看应用启动完成的日志,以此调整initialDelaySeconds的准确数值。

2. 调整滚动更新策略

为严格实现零停机,修改Deployment的滚动更新配置,确保旧Pod仅在新Pod完全就绪后才会被销毁:

strategy:
  rollingUpdate:
    maxSurge: 1  # 允许额外启动1个Pod(适合副本数较少的场景)
    maxUnavailable: 0  # 不允许任何Pod不可用,保证旧Pod持续运行至新Pod就绪
  type: RollingUpdate

此配置下,Kubernetes会先启动新Pod,待新Pod通过就绪探针后,再销毁旧Pod,彻底避免服务中断。

3. 验证Actuator端点可用性

确保Spring Boot应用正确暴露了健康探针端点:

  • 检查application.properties中是否开启了Actuator的Web访问:
# 暴露健康相关端点
management.endpoints.web.exposure.include=health
# 或直接暴露所有端点
management.endpoints.web.exposure.include=*
  • 本地启动应用后,访问http://localhost:8000/actuator/health/liveness和http://localhost:8000/actuator/health/readiness,确认返回UP状态。

验证步骤

  1. 应用修改后的配置:kubectl apply -f deployment.yaml
  2. 实时查看Pod状态:kubectl get pods -w,确认新Pod成功进入Running状态且就绪状态为1/1
  3. 部署过程中测试currency-conversion服务,确认无内部服务器错误

内容的提问来源于stack exchange,提问作者Haitam-Elgharras

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 14:22:48