You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Docker环境下Supervisor管理的Gunicorn服务频繁停止问题

问题描述

我通过Docker结合Supervisor、Nginx和Gunicorn部署Django项目,用docker-compose up命令构建项目及Nginx镜像。查同类问题时看到建议用/home/ubuntu/...这类绝对路径,但我在Docker容器里通过COPY . /app/把项目复制到了/app目录,不清楚Supervisor配置文件里的directory字段该怎么填。

Gunicorn已通过requirements.txt安装完成,我的Supervisor配置如下:

[supervisord]
nodaemon=true

[program:gunicorn]
command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000
directory=/app
user=root
autostart=true
autorestart=true
stderr_logfile=/var/log/gunicorn.err.log
stdout_logfile=/var/log/gunicorn.out.log

[program:nginx]
command=nginx -g "daemon off;"
autostart=true
autorestart=true
stderr_logfile=/var/log/nginx.err.log
stdout_logfile=/var/log/nginx.out.log

错误日志

docker-compose logs -f web输出:

web-1  | 2024-06-26 21:41:45,768 CRIT Supervisor is running as root.  Privileges were not dropped because no user is specified in the config file.  If you intend to run as root, you can set user=root in the config file to avoid this message.
web-1  | 2024-06-26 21:41:45,771 INFO supervisord started with pid 1
web-1  | 2024-06-26 21:41:46,775 INFO spawned: 'gunicorn' with pid 7
web-1  | 2024-06-26 21:41:46,779 INFO spawned: 'nginx' with pid 8
web-1  | 2024-06-26 21:41:47,365 WARN exited: gunicorn (exit status 3; not expected)
web-1  | 2024-06-26 21:41:48,368 INFO spawned: 'gunicorn' with pid 11
web-1  | 2024-06-26 21:41:48,368 INFO success: nginx entered RUNNING state, process has stayed up for > than 1 seconds (startsecs)
web-1  | 2024-06-26 21:41:50,031 INFO success: gunicorn entered RUNNING state, process has stayed up for > than 1 seconds (startsecs)
web-1  | 2024-06-26 21:41:50,031 WARN exited: gunicorn (exit status 1; not expected)
web-1  | 2024-06-26 21:41:51,034 INFO spawned: 'gunicorn' with pid 15
web-1  | 2024-06-26 21:41:51,612 WARN exited: gunicorn (exit status 3; not expected)
web-1  | 2024-06-26 21:41:52,615 INFO spawned: 'gunicorn' with pid 17
web-1  | 2024-06-26 21:41:53,181 WARN exited: gunicorn (exit status 3; not expected)
web-1  | 2024-06-26 21:41:55,185 INFO spawned: 'gunicorn' with pid 19
web-1  | 2024-06-26 21:41:55,753 WARN exited: gunicorn (exit status 3; not expected)
web-1  | 2024-06-26 21:41:58,758 INFO spawned: 'gunicorn' with pid 21
web-1  | 2024-06-26 21:41:59,320 WARN exited: gunicorn (exit status 3; not expected)
web-1  | 2024-06-26 21:42:00,321 INFO gave up: gunicorn entered FATAL state, too many start retries too quickly

gunicorn.err.log中的具体错误:

Traceback (most recent call last):
  File "/usr/local/bin/gunicorn", line 8, in <module>
    sys.exit(run())
             ^^^^^
  File "/usr/local/lib/python3.11/site-packages/gunicorn/app/wsgiapp.py", line 67, in run
    WSGIApplication("%(prog)s [OPTIONS] [APP_MODULE]", prog=prog).run()
  File "/usr/local/lib/python3.11/site-packages/gunicorn/app/base.py", line 236, in run
    super().run()
  File "/usr/local/lib/python3.11/site-packages/gunicorn/app/base.py", line 72, in run
    Arbiter(self).run()
  File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 229, in run
    self.halt(reason=inst.reason, exit_status=inst.exit_status)
  File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 342, in halt
    self.stop()
  File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 396, in stop
    time.sleep(0.1)
  File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 242, in handle_chld
    self.reap_workers()
  File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 530, in reap_workers
    raise HaltServer(reason, self.WORKER_BOOT_ERROR)
gunicorn.errors.HaltServer: <HaltServer 'Worker failed to boot.' 3>

我的配置文件

Dockerfile

FROM python:3.11.9-slim

ENV PYTHONDONTWRITEBYTECODE 1
ENV PYTHONUNBUFFERED 1

WORKDIR /app

RUN apt-get update \
    && apt-get install -y --no-install-recommends \
        gcc \
        python3-dev \
        musl-dev \
        postgresql-client \
        libpq-dev \
        nginx \
        supervisor \
    && rm -rf /var/lib/apt/lists/*

RUN rm /etc/nginx/nginx.conf
COPY nginx.conf /etc/nginx/nginx.conf

COPY requirements.txt /app/
RUN pip install --no-cache-dir -r requirements.txt

COPY . /app/

RUN python manage.py collectstatic --no-input
RUN python manage.py makemigrations
RUN python manage.py migrate

COPY supervisord.conf /etc/supervisor/conf.d/supervisord.conf

EXPOSE 80

CMD ["/usr/bin/supervisord", "-c", "/etc/supervisor/conf.d/supervisord.conf"]

docker-compose.yml

services:
  web:
    build: .
    command: /usr/bin/supervisord -c /etc/supervisor/conf.d/supervisord.conf
    environment:
      DATABASE_NAME: ###
      DATABASE_USER: ###
      DATABASE_PASSWORD: ###
      DATABASE_HOST: ###
      DATABASE_PORT: 5432
    volumes:
      - /path_to_static:/app/static #edited for stackoverflow
      - /path_to_static:/app/media #edited for stackoverflow
    networks:
      - backend

  nginx:
    build:
      context: .
      dockerfile: Dockerfile.nginx
    ports:
      - "80:80"
    volumes:
      - /path_to_static:/app/static #edited for stackoverflow
      - /path_to_static:/app/media #edited for stackoverflow
    networks:
      - backend
    depends_on:
      - web

networks:
  backend:

Dockerfile.nginx

FROM nginx:latest
RUN rm /etc/nginx/nginx.conf
COPY nginx.conf /etc/nginx/nginx.conf

解决步骤

1. 确认Supervisor的directory字段配置

你的directory=/app是正确的——项目已复制到容器的/app目录,且Dockerfile中设置了WORKDIR /app,这部分无需修改。

2. 定位Gunicorn启动失败的具体原因

Gunicorn退出码3表示Worker启动失败,但当前日志未展示Django层面的报错。修改Supervisor中gunicorn的命令,开启调试日志:

command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000 --log-level debug

重新启动后,查看gunicorn.err.log就能看到具体错误(比如Django配置错误、数据库连接失败等)。

3. 修复Dockerfile中的核心问题

  • 迁移时机错误:Docker build阶段无法访问docker-compose网络中的数据库,因此不能在build时执行makemigrations和migrate。移除Dockerfile中的这两行:
    RUN python manage.py makemigrations
    RUN python manage.py migrate
    
    改用entrypoint脚本在容器启动时执行迁移:
    创建entrypoint.sh文件:
    #!/bin/sh
    python manage.py makemigrations
    python manage.py migrate
    python manage.py collectstatic --no-input
    exec "$@"
    
    在Dockerfile中添加:
    COPY entrypoint.sh /app/
    RUN chmod +x /app/entrypoint.sh
    ENTRYPOINT ["/app/entrypoint.sh"]
    
  • 重复部署Nginx:web容器和docker-compose中各部署了一个Nginx,会导致资源冲突。保留独立的Nginx服务,移除web容器中的Nginx相关配置:
    修改Dockerfile的安装命令,去掉nginx:
    RUN apt-get update \
        && apt-get install -y --no-install-recommends \
            gcc \
            python3-dev \
            musl-dev \
            postgresql-client \
            libpq-dev \
            supervisor \
        && rm -rf /var/lib/apt/lists/*
    
    同时删除Supervisor配置中的[program:nginx]块。

4. 验证数据库连接

确认docker-compose中的数据库环境变量正确,且数据库服务在同一个backend网络中。可以在容器启动后手动测试连接:

docker-compose exec web psql -h $DATABASE_HOST -U $DATABASE_USER -d $DATABASE_NAME

5. 优化Supervisor配置

  • 提前创建日志目录,避免启动报错:
    在Dockerfile中添加:
    RUN mkdir -p /var/log/gunicorn
    
  • 为Gunicorn添加稳定性参数:
    command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000 --workers 2 --timeout 30 --log-level debug
    

内容的提问来源于stack exchange,提问作者Uday Singh Khalsa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.21 18:42:34