You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Azure Container App健康探测报错排查求助

问题概述

我有一个运行Node.js和Redis服务的API应用,仅暴露3000端口,对接PostgreSQL进行数据查询与存储,所有服务都运行在单个容器中。已配置存活(liveness)、就绪(readiness)和启动(startup)探测,但系统日志仍持续出现这三类探测报错。当前已启用HTTP类型的Ingress,应用每日接收约15万请求,希望明确问题原因及解决方案。


当前配置

健康探测配置

已配置启动、就绪、存活三类HTTP探测,但持续报错(原配置含探测路径、超时、周期等参数)。

Dockerfile

FROM node:21.4-bookworm

RUN apt-get update && \
    apt-get install -y iputils-ping traceroute telnet dnsutils git nano htop redis-server && \
    npm install -g pm2 && \
    ln -fs /usr/share/zoneinfo/America/Santiago /etc/localtime && \
    rm -rf /var/lib/apt/lists/*

RUN echo "user1:password" | chpasswd

RUN adduser -u 2000 user2 && \
    echo "user2:password" | chpasswd

WORKDIR /usr/src/app/CUA-Tunel

COPY . .
    
EXPOSE 3000

COPY entrypoint.sh /entrypoint.sh

RUN chmod +x /entrypoint.sh

ENTRYPOINT ["/entrypoint.sh"]

entrypoint.sh

#!/bin/bash

# Configuración DNS
echo search server01.domain.com > /etc/resolv.conf
echo search server02.domain.com >> /etc/resolv.conf
echo "nameserver xxx.xxx.xxx.xxx" >> /etc/resolv.conf
echo "nameserver xxx.xxx.xxx.xxx" >> /etc/resolv.conf

echo "maxclients 500000" >> /etc/redis/redis.conf

npm install

/etc/init.d/redis-server start

npx pm2 start ecosystem.config.js --no-daemon

docker-compose.yml

#version: '2.0'

services:
  tunel-nodejs:
    build:
      context: .
      dockerfile: Dockerfile
    ports:
      - "3000:3000"

问题原因分析

  1. 单容器多进程依赖冲突:容器内同时运行Redis、Node.js两个服务,启动脚本串行执行但未等待Redis就绪,Node服务启动时可能因Redis未初始化完成而无法正常响应;高请求量下,Redis或Node进程占用过多资源会互相影响,导致服务无法响应探测。
  2. 启动流程耗时过长:容器启动时执行npm install,这个步骤会大幅增加启动时间,若启动探测的初始等待时间过短,会提前触发探测导致失败。
  3. DNS配置错误:直接覆盖/etc/resolv.conf可能破坏容器默认DNS解析,多个search域会增加DNS解析延迟,导致Node服务无法连接PostgreSQL,进而无法响应探测请求。
  4. 探测参数不匹配:启动探测的等待时间、周期设置过短,未给服务足够的初始化时间;存活/就绪探测的超时、重试次数设置不合理,高负载下服务偶尔响应慢就会触发探测失败。
  5. 权限与进程管理隐患:以root用户启动服务,但创建了非root用户user2,可能存在文件/端口权限问题;PM2以no-daemon模式运行,若Node进程崩溃重启,期间探测会失败。

解决方案

1. 拆分服务为独立容器

将Redis、Node.js拆分为单独容器,避免单容器多进程的互相影响,便于单独监控和健康管理。修改docker-compose.yml示例:

version: '3.8'

services:
  redis:
    image: redis:latest
    command: redis-server --maxclients 500000
    restart: always
  tunel-nodejs:
    build:
      context: .
      dockerfile: Dockerfile
    ports:
      - "3000:3000"
    depends_on:
      - redis
    dns:
      - xxx.xxx.xxx.xxx
      - xxx.xxx.xxx.xxx
    dns_search:
      - server01.domain.com
      - server02.domain.com
    restart: always

同时修改Dockerfile,移除Redis安装步骤,entrypoint中不再启动Redis。

2. 优化启动流程

  • 将npm install移至Dockerfile构建阶段,减少容器启动时间:
    在Dockerfile的COPY . .后添加:
    RUN npm install
    
    并删除entrypoint中的npm install命令。
  • 若未拆分容器,在entrypoint中添加Redis就绪等待逻辑:
    # 等待Redis就绪
    until redis-cli ping; do
        echo "等待Redis启动..."
        sleep 2
    done
    
    放在启动PM2之前。

3. 调整健康探测参数

根据服务实际启动时间和响应能力调整探测参数(以Kubernetes为例):

livenessProbe:
  httpGet:
    path: /health
    port: 3000
  initialDelaySeconds: 30
  timeoutSeconds: 3
  periodSeconds: 5
  failureThreshold: 3
readinessProbe:
  httpGet:
    path: /health
    port: 3000
  initialDelaySeconds: 10
  timeoutSeconds: 2
  periodSeconds: 3
  failureThreshold: 2
startupProbe:
  httpGet:
    path: /health
    port: 3000
  initialDelaySeconds: 60
  periodSeconds: 10
  failureThreshold: 10

确保/health接口是轻量级的,仅检查Redis、PostgreSQL连接状态,不执行复杂业务逻辑。

4. 修复DNS配置

删除entrypoint中修改/etc/resolv.conf的代码,改用Docker Compose或Kubernetes的DNS配置参数(如上述compose示例中的dns和dns_search),避免破坏容器默认DNS解析。

5. 优化进程与权限管理

  • 在Dockerfile末尾添加USER user2,以非root用户运行服务,提前调整文件权限:
    RUN chown -R user2:user2 /usr/src/app/CUA-Tunel
    USER user2
    
  • 优化PM2配置文件ecosystem.config.js,添加进程监控和重启策略,确保Node进程异常时快速恢复。

6. 日志与监控排查

  • 确保Node服务和Redis的日志输出到stdout/stderr,便于通过docker logs或Kubernetes日志工具查看探测失败时的具体错误。
  • 监控容器资源使用情况(CPU、内存),排查是否因资源不足导致服务无法响应探测。

内容的提问来源于stack exchange,提问作者Cristopher

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 02:55:11