排查K8s环境下Django应用Readiness Probe失败问题求助
Django REST Framework应用Kubernetes就绪探针失败排查
问题概述
在Kubernetes环境中部署Django REST Framework应用时,就绪探针(Readiness Probe)检测失败。Pod日志显示应用已正常启动,仅存在1个未应用的auth模块迁移(暂不影响运行),但探针返回连接拒绝错误。
Pod日志输出
C:\Users\user>kubectl logs test-bbdccbc76-8cwg9 Watching for file changes with StatReloader Performing system checks... Server initialized for gevent. System check identified some issues: WARNINGS: ?: (staticfiles.W004) The directory '/app/staticfiles' in the STATICFILES_DIRS setting does not exist. System check identified 1 issue (0 silenced). You have 1 unapplied migration(s). Your project may not work properly until you apply the migrations for app(s): auth Run 'python manage.py migrate' to apply them. December 27, 2022 - 06:13:28 Django version 4.1.2, using settings 'test.settings' Starting development server at http://127.0.0.1:8000/ Quit the server with CONTROL-C.
Pod描述输出
C:\Users\user>kubectl describe pod test-bbdccbc76-8cwg9 Name: test-bbdccbc76-8cwg9 Namespace: test Priority: 2000 Priority Class Name: default Service Account: default Node: sd2-k8s-stg-n05/10.216.14.52 Start Time: Tue, 27 Dec 2022 11:11:14 +0500 Labels: app=test pod-template-hash=bbdccbc76 Annotations: checksum/config: f0ee0887d3fd078979831f04d13ade8759ef8a4ee9aad23830c5909300e322b4 cni.projectcalico.org/containerID: ac4c0c57c0a3baa1c15a78aedb36e3e1890f8b9859547807b1e3587835792efb cni.projectcalico.org/podIP: 10.233.110.84/32 cni.projectcalico.org/podIPs: 10.233.110.84/32 container.apparmor.security.beta.kubernetes.io/test: runtime/default kubernetes.io/psp: restricted seccomp.security.alpha.kubernetes.io/pod: runtime/default Status: Running IP: 10.233.110.84 IPs: IP: 10.233.110.84 Controlled By: ReplicaSet/test-bbdccbc76 Containers: test: Container ID: containerd://26e3234ba31c09d6d62e3efb999634cb64d26d6ce2e8e734a26104d62b6f2f5f Image: registry/library/test:latest Image ID: registry/library/test-@sha256:12345678b16eb2c90324916756846d5dfa557198bd4aeed9e790db677702b1 Port: 8000/TCP Host Port: 0/TCP Command: python manage.py runserver State: Running Started: Tue, 27 Dec 2022 11:11:15 +0500 Ready: False Restart Count: 0 Limits: cpu: 500m ephemeral-storage: 500Mi memory: 500Mi Requests: cpu: 500m ephemeral-storage: 500Mi memory: 500M Liveness: http-get http://:8000/healthz delay=30s timeout=1s period=10s #success=1 #failure=3 Readiness: http-get http://:8000/healthz delay=30s timeout=1s period=10s #success=1 #failure=3 Environment Variables from: test-configmap ConfigMap Optional: false Environment: CONNECTION_STRING: mongodb://xxx:yyy@zzz:27017 Mounts: /var/run/secrets/kubernetes.io/serviceaccount from kube-api-access-tpstr (ro) Conditions: Type Status Initialized True Ready False ContainersReady False PodScheduled True Volumes: kube-api-access-tpstr: Type: Projected (a volume that contains injected data from multiple sources) TokenExpirationSeconds: 3607 ConfigMapName: kube-root-ca.crt ConfigMapOptional: <nil> DownwardAPI: true QoS Class: Burstable Node-Selectors: <none> Tolerations: node.kubernetes.io/not-ready:NoExecute op=Exists for 300s node.kubernetes.io/unreachable:NoExecute op=Exists for 300s Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 79s default-scheduler Successfully assigned test/test-bbdccbc76-8cwg9 to xx-k8s-xx Normal Pulling 78s kubelet Pulling image "registry/library/test:latest" Normal Pulled 78s kubelet Successfully pulled image "registry/library/test:latest" in 57.066835ms Normal Created 78s kubelet Created container test Normal Started 78s kubelet Started container test Warning Unhealthy 29s (x2 over 39s) kubelet Readiness probe failed: Get "http://10.233.110.84:8000/healthz": dial tcp 10.233.110.84:8000: connect: connection refused Warning Unhealthy 29s (x2 over 39s) kubelet Liveness probe failed: Get "http://10.233.110.84:8000/healthz": dial tcp 10.233.110.84:8000: connect: connection refused
Service描述输出
C:\Users\user>kubectl describe service test-service Name: test-service Namespace: test Labels: app=test app.kubernetes.io/managed-by=Helm app.kubernetes.io/version=0.1.0 helm.sh/chart=my-helm-char-1.0.0 Annotations: meta.helm.sh/release-name: test meta.helm.sh/release-namespace: test Selector: app=test Type: NodePort IP Family Policy: SingleStack IP Families: IPv4 IP: 10.233.53.185 IPs: 10.233.53.185 Port: api-port 8000/TCP TargetPort: api-port/TCP NodePort: api-port 32426/TCP Endpoints: Session Affinity: None External Traffic Policy: Cluster Events: <none>
问题排查与解决步骤
1. 核心问题:Django服务器绑定地址限制
从Pod日志可见,Django开发服务器启动时绑定的是127.0.0.1:8000,仅监听容器内部的回环接口,而Kubernetes探针尝试访问Pod的IP(10.233.110.84),因此无法建立连接,导致探针失败。
解决方法:修改容器启动命令,让服务器绑定所有网卡:
python manage.py runserver 0.0.0.0:8000
在Deployment的容器配置中更新Command字段即可。
2. 验证健康检查端点存在性
Kubernetes探针配置的是/healthz路径,但Django默认没有提供该端点。需确认:
- 若使用第三方库(如
django-health-check),需确保已正确安装并配置路由 - 若自定义健康检查视图,需将其映射到
/healthz路径
3. 处理未应用迁移
虽然当前不影响探针检测,但未应用的auth模块迁移可能导致后续功能异常,建议在容器启动时添加迁移命令,或在镜像构建阶段执行:
python manage.py migrate
4. 解决静态文件目录警告
日志提示/app/staticfiles目录不存在,可通过以下方式处理:
- 在Dockerfile中添加创建目录的命令:
RUN mkdir -p /app/staticfiles - 修改Django配置文件中的
STATICFILES_DIRS,指向存在的目录
内容的提问来源于stack exchange,提问作者StuffHappens
相关产品推荐
相关产品推荐

