Google Cloud Run容器健康检查内部错误致部署失败求助
问题场景
部署新版本时偶发失败,无特定触发因素,错误提示为**"执行容器健康检查时发生内部错误"**,部署命令为:
gcloud alpha run services replace ...
服务配置(cloudrun.yml)
apiVersion: serving.knative.dev/v1 kind: Service metadata: name: <my_app> spec: traffic: - percent: 100 latestRevision: true template: metadata: annotations: autoscaling.knative.dev/minScale: '1' autoscaling.knative.dev/maxScale: '40' run.googleapis.com/vpc-access-connector: <my_app>-vpc-16gbps run.googleapis.com/sandbox: gvisor spec: timeoutSeconds: 50 serviceAccountName: ... containerConcurrency: 100 containers: - image: ... ports: - containerPort: 8080 name: h2c resources: limits: cpu: '4' memory: 6Gi
审计日志错误信息
{ "protoPayload": { "@type": "type.googleapis.com/google.cloud.audit.AuditLog", "status": { "code": 13, "message": "Ready condition status changed to False for Service <my_app> with message: Internal error occurred while performing container health check. Resource readiness deadline exceeded." }, "serviceName": "run.googleapis.com", "response": { "apiVersion": "serving.knative.dev/v1", "kind": "Service", "status": { "observedGeneration": 1, "conditions": [ { "type": "Ready", "status": "False", "message": "Internal error occurred while performing container health check. Resource readiness deadline exceeded.", "lastTransitionTime": "2022-10-05T14:11:19.884568Z" }, { "type": "ConfigurationsReady", "status": "Unknown", "message": "Internal error occurred while performing container health check.", "lastTransitionTime": "2022-10-05T14:00:50.398740Z" }, { "type": "RoutesReady", "status": "False", "reason": "RevisionFailed", "message": "Revision '<my_app>-s7px8' is not ready and cannot serve traffic. Internal error occurred while performing container health check.", "lastTransitionTime": "2022-10-05T14:11:19.884568Z" } ], "latestCreatedRevisionName": "<my_app>-s7px8", }, "@type": "type.googleapis.com/google.cloud.run.v1.Service" } } }
关键信息:应用日志无异常,部署15-20秒即失败(远早于默认TCP探针4分钟超时),未配置任何自定义启动/健康/存活探针。
排查方向建议
排查VPC访问连接器的稳定性
配置的VPC连接器可能存在临时资源瓶颈或故障,导致健康检查流量无法正常转发。查看连接器的监控指标(CPU、内存、并发连接数),确认配额是否满足当前需求,必要时调整连接器的规格。验证gVisor沙箱的兼容性
gVisor沙箱的网络层可能存在偶发兼容性问题,尝试临时移除run.googleapis.com/sandbox: gvisor注解后重新部署,观察是否还会出现失败,以此排除沙箱的影响。检查部署区域的资源负载
偶发失败可能和部署区域的临时资源紧张有关,查看GCP控制台中该区域的Cloud Run服务状态,或尝试切换到同区域内的其他可用区(如果支持)测试。自定义启动探针延长超时窗口
虽然默认TCP探针超时为4分钟,但实际就绪deadline可能受其他因素压缩。可以自定义启动探针,设置更长的等待和重试周期,示例配置:containers: - image: ... ports: - containerPort: 8080 name: h2c startupProbe: tcpSocket: port: 8080 initialDelaySeconds: 10 timeoutSeconds: 5 periodSeconds: 5 failureThreshold: 24此配置下总超时时间可达130秒,给容器足够的启动缓冲时间。
确认容器启动时的端口监听时机
在容器启动脚本中添加日志,记录应用程序完成端口监听的准确时间,确认是否在Cloud Run发起健康检查前,应用已经完成端口初始化。查看GCP服务端状态
偶发内部错误可能是GCP服务端的临时故障,检查GCP控制台的服务状态页面,确认部署期间Cloud Run是否有已知的异常事件。
内容的提问来源于stack exchange,提问作者star67

