You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

排查K8s环境下Django应用Readiness Probe失败问题求助

Django REST Framework应用Kubernetes就绪探针失败排查

问题概述

在Kubernetes环境中部署Django REST Framework应用时,就绪探针(Readiness Probe)检测失败。Pod日志显示应用已正常启动,仅存在1个未应用的auth模块迁移(暂不影响运行),但探针返回连接拒绝错误。

Pod日志输出

C:\Users\user>kubectl logs test-bbdccbc76-8cwg9
Watching for file changes with StatReloader
Performing system checks...

Server initialized for gevent.
System check identified some issues:

WARNINGS:
?: (staticfiles.W004) The directory '/app/staticfiles' in the STATICFILES_DIRS setting does not exist.

System check identified 1 issue (0 silenced).

You have 1 unapplied migration(s). Your project may not work properly until you apply the migrations for app(s): auth
Run 'python manage.py migrate' to apply them.
December 27, 2022 - 06:13:28
Django version 4.1.2, using settings 'test.settings'
Starting development server at http://127.0.0.1:8000/
Quit the server with CONTROL-C.

Pod描述输出

C:\Users\user>kubectl describe pod test-bbdccbc76-8cwg9
Name:                 test-bbdccbc76-8cwg9
Namespace:            test
Priority:             2000
Priority Class Name:  default
Service Account:      default
Node:                 sd2-k8s-stg-n05/10.216.14.52
Start Time:           Tue, 27 Dec 2022 11:11:14 +0500
Labels:               app=test
                      pod-template-hash=bbdccbc76
Annotations:          checksum/config: f0ee0887d3fd078979831f04d13ade8759ef8a4ee9aad23830c5909300e322b4
                      cni.projectcalico.org/containerID: ac4c0c57c0a3baa1c15a78aedb36e3e1890f8b9859547807b1e3587835792efb
                      cni.projectcalico.org/podIP: 10.233.110.84/32
                      cni.projectcalico.org/podIPs: 10.233.110.84/32
                      container.apparmor.security.beta.kubernetes.io/test: runtime/default
                      kubernetes.io/psp: restricted
                      seccomp.security.alpha.kubernetes.io/pod: runtime/default
Status:               Running
IP:                   10.233.110.84
IPs:
  IP:           10.233.110.84
Controlled By:  ReplicaSet/test-bbdccbc76
Containers:
  test:
    Container ID:  containerd://26e3234ba31c09d6d62e3efb999634cb64d26d6ce2e8e734a26104d62b6f2f5f
    Image:         registry/library/test:latest
    Image ID:      registry/library/test-@sha256:12345678b16eb2c90324916756846d5dfa557198bd4aeed9e790db677702b1
    Port:          8000/TCP
    Host Port:     0/TCP
    Command:
      python
      manage.py
      runserver
    State:          Running
      Started:      Tue, 27 Dec 2022 11:11:15 +0500
    Ready:          False
    Restart Count:  0
    Limits:
      cpu:                500m
      ephemeral-storage:  500Mi
      memory:             500Mi
    Requests:
      cpu:                500m
      ephemeral-storage:  500Mi
      memory:             500M
    Liveness:             http-get http://:8000/healthz delay=30s timeout=1s period=10s #success=1 #failure=3
    Readiness:            http-get http://:8000/healthz delay=30s timeout=1s period=10s #success=1 #failure=3
    Environment Variables from:
      test-configmap  ConfigMap  Optional: false
    Environment:
      CONNECTION_STRING:     mongodb://xxx:yyy@zzz:27017
    Mounts:
      /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-tpstr (ro)
Conditions:
  Type              Status
  Initialized       True
  Ready             False
  ContainersReady   False
  PodScheduled      True
Volumes:
  kube-api-access-tpstr:
    Type:                    Projected (a volume that contains injected data from multiple sources)
    TokenExpirationSeconds:  3607
    ConfigMapName:           kube-root-ca.crt
    ConfigMapOptional:       <nil>
    DownwardAPI:             true
QoS Class:                   Burstable
Node-Selectors:              <none>
Tolerations:                 node.kubernetes.io/not-ready:NoExecute op=Exists for 300s
                             node.kubernetes.io/unreachable:NoExecute op=Exists for 300s

Events:
  Type     Reason     Age                From               Message
  ----     ------     ----               ----               -------
  Normal   Scheduled  79s                default-scheduler  Successfully assigned test/test-bbdccbc76-8cwg9 to xx-k8s-xx
  Normal   Pulling    78s                kubelet            Pulling image "registry/library/test:latest"
  Normal   Pulled     78s                kubelet            Successfully pulled image "registry/library/test:latest" in 57.066835ms
  Normal   Created    78s                kubelet            Created container test
  Normal   Started    78s                kubelet            Started container test
  Warning  Unhealthy  29s (x2 over 39s)  kubelet            Readiness probe failed: Get "http://10.233.110.84:8000/healthz": dial tcp 10.233.110.84:8000: connect: connection refused
  Warning  Unhealthy  29s (x2 over 39s)  kubelet            Liveness probe failed: Get "http://10.233.110.84:8000/healthz": dial tcp 10.233.110.84:8000: connect: connection refused

Service描述输出

C:\Users\user>kubectl describe service test-service
Name:                     test-service
Namespace:                test
Labels:                   app=test
                          app.kubernetes.io/managed-by=Helm
                          app.kubernetes.io/version=0.1.0
                          helm.sh/chart=my-helm-char-1.0.0
Annotations:              meta.helm.sh/release-name: test
                          meta.helm.sh/release-namespace: test
Selector:                 app=test
Type:                     NodePort
IP Family Policy:         SingleStack
IP Families:              IPv4
IP:                       10.233.53.185
IPs:                      10.233.53.185
Port:                     api-port  8000/TCP
TargetPort:               api-port/TCP
NodePort:                 api-port  32426/TCP
Endpoints:
Session Affinity:         None
External Traffic Policy:  Cluster
Events:                   <none>

问题排查与解决步骤

1. 核心问题:Django服务器绑定地址限制

从Pod日志可见,Django开发服务器启动时绑定的是127.0.0.1:8000,仅监听容器内部的回环接口,而Kubernetes探针尝试访问Pod的IP(10.233.110.84),因此无法建立连接,导致探针失败。

解决方法:修改容器启动命令,让服务器绑定所有网卡:

python manage.py runserver 0.0.0.0:8000

在Deployment的容器配置中更新Command字段即可。

2. 验证健康检查端点存在性

Kubernetes探针配置的是/healthz路径,但Django默认没有提供该端点。需确认:

  • 若使用第三方库(如django-health-check),需确保已正确安装并配置路由
  • 若自定义健康检查视图,需将其映射到/healthz路径

3. 处理未应用迁移

虽然当前不影响探针检测,但未应用的auth模块迁移可能导致后续功能异常,建议在容器启动时添加迁移命令,或在镜像构建阶段执行:

python manage.py migrate

4. 解决静态文件目录警告

日志提示/app/staticfiles目录不存在,可通过以下方式处理:

  • 在Dockerfile中添加创建目录的命令:RUN mkdir -p /app/staticfiles
  • 修改Django配置文件中的STATICFILES_DIRS,指向存在的目录

内容的提问来源于stack exchange,提问作者StuffHappens

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 14:00:33