You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes单Pod多容器场景下单容器崩溃时如何重启整个Pod

实现方案

Kubernetes默认容器重启逻辑为单容器独立重启,要实现任意容器崩溃触发整Pod重启,可以通过以下两种方式实现:

方案1:Kubernetes 1.29+ 原生特性实现

Kubernetes 1.29及以上版本稳定了PodContainerRestartPolicy特性,直接配置即可原生支持全容器重启逻辑,无需额外自定义组件。
配置示例:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: multi-container-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: multi-container-demo
  template:
    metadata:
      labels:
        app: multi-container-demo
    spec:
      restartPolicy: Always
      # 核心配置:指定所有容器异常时触发整Pod重启
      containerRestartPolicy: AllContainers
      containers:
      - name: container-1
        image: nginx:alpine
      - name: container-2
        image: busybox:1.36
        command: ["sleep", "3600"]
      - name: container-3
        image: redis:7-alpine
      - name: container-4
        image: postgres:15-alpine

注意:使用该特性需要确认集群所有组件(kube-apiserver、kubelet)已开启PodContainerRestartPolicy特性门控


方案2:全版本兼容通用方案

如果集群版本低于1.29,可以通过共享进程命名空间+哨兵容器的方式实现,适配所有Kubernetes版本。
实现逻辑:

  • 开启Pod的shareProcessNamespace配置,允许同Pod内的容器互相可见进程
  • 新增一个哨兵容器,定期检查所有业务容器的主进程存活状态
  • 只要任意业务容器主进程退出,哨兵容器主动异常退出,触发Pod健康检查失败,进而重建整个Pod
    配置示例片段:
apiVersion: apps/v1
kind: Deployment
metadata:
  name: multi-container-demo
spec:
  replicas: 1
  selector:
    matchLabels:
      app: multi-container-demo
  template:
    metadata:
      labels:
        app: multi-container-demo
    spec:
      # 开启进程共享
      shareProcessNamespace: true
      containers:
      # 你的4个业务容器配置,此处省略
      - name: container-1
        image: nginx:alpine
      - name: container-2
        image: busybox:1.36
        command: ["sleep", "3600"]
      - name: container-3
        image: redis:7-alpine
      - name: container-4
        image: postgres:15-alpine
      # 新增哨兵容器
      - name: crash-sentinel
        image: alpine:3.18
        securityContext:
          privileged: true
        command: ["/bin/sh", "-c"]
        args:
        - |
          # 替换为你4个业务容器的主进程关键词,确保唯一匹配
          PROCESS_LIST=("nginx" "sleep" "redis-server" "postgres")
          while true; do
            for proc in "${PROCESS_LIST[@]}"; do
              if ! pgrep -f "$proc" > /dev/null; then
                echo "进程 $proc 已退出,触发整Pod重启"
                exit 1
              fi
            done
            sleep 2
          done
        livenessProbe:
          exec:
            command: ["pgrep", "-f", "crash-sentinel"]
          initialDelaySeconds: 5
          periodSeconds: 2

注意事项

  • 哨兵容器的进程关键词需要和业务容器主进程完全匹配,避免误判或者漏判
  • 若业务容器为多进程架构,可调整检查逻辑为检查对应容器的PID 1进程状态
  • 若使用原生特性方案,需确认集群版本和特性门控配置符合要求

内容的提问来源于stack exchange,提问作者Kufu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 08:15:03