You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Kubernetes中Python脚本Pod遇CrashLoopBackOff问题及定时执行需求

Kubernetes定时运行Python脚本的CrashLoopBackOff问题排查

需求背景

需在Kubernetes中定时运行生成文件的Python脚本,最终通过CronJob实现并将结果保存至关联本地存储。先尝试用Docker镜像一次性运行脚本,但Pod出现CrashLoopBackOff状态。

相关代码与配置

Python脚本(create_txt.py)

import numpy as np
import datetime

res = np.random.rand(1)[0]
res = np.round(res,3) * 1000

with open(f'/home/sjw/kube/{str(int(res))}.txt','w') as f:
    txt = datetime.datetime.now().strftime("%H:%M:%S")
    f.write(txt)

Dockerfile配置

FROM python:3
WORKDIR /home/sjw/kube
COPY create_txt.py ./
RUN pip install numpy
CMD ["python","./create_txt.py"]

Pod清单

apiVersion: v1
kind: Pod
metadata:
  name: createinterval
spec:
  containers:
  - name: createinterval
    image: idioluck/kube_create:v01
    command: ["/bin/sh"]
    args: ["python create_txt.py"]
    volumeMounts:
    - mountPath: /home/sjw/kube
      name: testvol
  volumes:
  - name: testvol
    hostPath:
      path: /home/sjw/kube
      type: DirectoryOrCreate

Pod状态信息

NAME              READY   STATUS             RESTARTS     AGE
createinterval    0/1     CrashLoopBackOff   7 (90s ago)   12m

kubectl describe pod 输出

sjw@DESKTOP-O6E7MND:~/kube/docker_sample/kube_create_txt_interval$ kubectl describe pod createinterval
Name:         createinterval
Namespace:    default
Priority:     0
Node:         minikube/192.168.49.2
Start Time:   Mon, 05 Sep 2022 16:42:24 +0900
Labels:       <none>
Annotations:  <none>
Status:       Running
IP:           172.17.0.4
IPs:
  IP:  172.17.0.4
Containers:
  createinterval:
    Container ID:   docker://89d2fd4597e445bfd11dace1e06ab325572d2e3072d14df9892b31ebbc7fa7d1
    Image:          idioluck/kube_create:v02
    Image ID:       docker-pullable://idioluck/kube_create@sha256:0868e3dc569c88641a3db05adbf2be9387609f9a0d184869ac939e80b93af5bb
    Port:           <none>
    Host Port:      <none>
    State:          Waiting
      Reason:       CrashLoopBackOff
    Last State:     Terminated
      Reason:       Completed
      Exit Code:    0
      Started:      Mon, 05 Sep 2022 17:08:35 +0900
      Finished:     Mon, 05 Sep 2022 17:08:35 +0900
    Ready:          False
    Restart Count:  10
    Environment:    <none>
    Mounts:
      /home/sjw/kube from testvol (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-5n28q (ro)
Conditions:
  Type              Status
  Initialized       True
  Ready             False
  ContainersReady   False
  PodScheduled      True
Volumes:
  testvol:
    Type:          HostPath (bare host directory volume)
    Path:          /home/sjw/kube
    HostPathType:  DirectoryOrCreate
  kube-api-access-5n28q:
    Type:                    Projected (a volume that contains injected data from multiple sources)
    TokenExpirationSeconds:  3607
    ConfigMapName:           kube-root-ca.crt
    ConfigMapOptional:       <nil>
    DownwardAPI:             true
QoS Class:                   BestEffort
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s
Events:
  Type     Reason     Age                   From               Message
  ----     ------     ----                  ----               -------
  Normal   Scheduled  27m                   default-scheduler  Successfully assigned default/createinterval to minikube
  Normal   Pulling    27m                   kubelet            Pulling image "idioluck/kube_create:v02"
  Normal   Pulled     26m                   kubelet            Successfully pulled image "idioluck/kube_create:v02" in 4.633892322s
  Normal   Created    25m (x5 over 26m)     kubelet            Created container createinterval
  Normal   Started    25m (x5 over 26m)     kubelet            Started container createinterval
  Normal   Pulled     25m (x4 over 26m)     kubelet            Container image "idioluck/kube_create:v02" already present on machine
  Warning  BackOff    109s (x117 over 26m)  kubelet            Back-off restarting failed container

问题分析与解决方案

从kubectl describe输出可见,容器退出码为0,说明脚本执行成功,但Kubernetes Pod默认会重启所有退出的容器,因此触发CrashLoopBackOff。

临时解决Pod循环重启问题

若仅测试脚本执行,无需Pod持续运行,可修改Pod的restartPolicy为OnFailure或Never:

apiVersion: v1
kind: Pod
metadata:
  name: createinterval
spec:
  restartPolicy: OnFailure  # 仅在容器非0退出码时重启
  containers:
  - name: createinterval
    image: idioluck/kube_create:v01
    command: ["python", "./create_txt.py"]  # 直接用Python执行,无需/bin/sh包裹
    volumeMounts:
    - mountPath: /home/sjw/kube
      name: testvol
  volumes:
  - name: testvol
    hostPath:
      path: /home/sjw/kube
      type: DirectoryOrCreate

最终实现:用CronJob定时运行

CronJob专为定时任务设计,执行完成后不会循环重启,符合需求。示例配置如下:

apiVersion: batch/v1
kind: CronJob
metadata:
  name: create-txt-job
spec:
  schedule: "*/5 * * * *"  # 每5分钟执行一次,可按需调整
  jobTemplate:
    spec:
      template:
        spec:
          containers:
          - name: create-txt
            image: idioluck/kube_create:v01
            command: ["python", "./create_txt.py"]
            volumeMounts:
            - mountPath: /home/sjw/kube
              name: testvol
          volumes:
          - name: testvol
            hostPath:
              path: /home/sjw/kube
              type: DirectoryOrCreate
          restartPolicy: OnFailure

额外优化点

  • 镜像瘦身:使用python:3-slim基础镜像减少体积,安装numpy时改用国内PyPI源加速。
  • 路径规范:避免依赖主机用户目录,改用Kubernetes PVC或标准存储路径,提升可移植性。
  • 日志增强:在Python脚本中添加日志输出,便于排查执行细节。

内容的提问来源于stack exchange,提问作者JaeWoo So

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 11:21:31