Docker环境下Supervisor管理的Gunicorn服务频繁停止问题
问题描述
我通过Docker结合Supervisor、Nginx和Gunicorn部署Django项目,用docker-compose up命令构建项目及Nginx镜像。查同类问题时看到建议用/home/ubuntu/...这类绝对路径,但我在Docker容器里通过COPY . /app/把项目复制到了/app目录,不清楚Supervisor配置文件里的directory字段该怎么填。
Gunicorn已通过requirements.txt安装完成,我的Supervisor配置如下:
[supervisord] nodaemon=true [program:gunicorn] command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000 directory=/app user=root autostart=true autorestart=true stderr_logfile=/var/log/gunicorn.err.log stdout_logfile=/var/log/gunicorn.out.log [program:nginx] command=nginx -g "daemon off;" autostart=true autorestart=true stderr_logfile=/var/log/nginx.err.log stdout_logfile=/var/log/nginx.out.log
错误日志
docker-compose logs -f web输出:
web-1 | 2024-06-26 21:41:45,768 CRIT Supervisor is running as root. Privileges were not dropped because no user is specified in the config file. If you intend to run as root, you can set user=root in the config file to avoid this message. web-1 | 2024-06-26 21:41:45,771 INFO supervisord started with pid 1 web-1 | 2024-06-26 21:41:46,775 INFO spawned: 'gunicorn' with pid 7 web-1 | 2024-06-26 21:41:46,779 INFO spawned: 'nginx' with pid 8 web-1 | 2024-06-26 21:41:47,365 WARN exited: gunicorn (exit status 3; not expected) web-1 | 2024-06-26 21:41:48,368 INFO spawned: 'gunicorn' with pid 11 web-1 | 2024-06-26 21:41:48,368 INFO success: nginx entered RUNNING state, process has stayed up for > than 1 seconds (startsecs) web-1 | 2024-06-26 21:41:50,031 INFO success: gunicorn entered RUNNING state, process has stayed up for > than 1 seconds (startsecs) web-1 | 2024-06-26 21:41:50,031 WARN exited: gunicorn (exit status 1; not expected) web-1 | 2024-06-26 21:41:51,034 INFO spawned: 'gunicorn' with pid 15 web-1 | 2024-06-26 21:41:51,612 WARN exited: gunicorn (exit status 3; not expected) web-1 | 2024-06-26 21:41:52,615 INFO spawned: 'gunicorn' with pid 17 web-1 | 2024-06-26 21:41:53,181 WARN exited: gunicorn (exit status 3; not expected) web-1 | 2024-06-26 21:41:55,185 INFO spawned: 'gunicorn' with pid 19 web-1 | 2024-06-26 21:41:55,753 WARN exited: gunicorn (exit status 3; not expected) web-1 | 2024-06-26 21:41:58,758 INFO spawned: 'gunicorn' with pid 21 web-1 | 2024-06-26 21:41:59,320 WARN exited: gunicorn (exit status 3; not expected) web-1 | 2024-06-26 21:42:00,321 INFO gave up: gunicorn entered FATAL state, too many start retries too quickly
gunicorn.err.log中的具体错误:
Traceback (most recent call last): File "/usr/local/bin/gunicorn", line 8, in <module> sys.exit(run()) ^^^^^ File "/usr/local/lib/python3.11/site-packages/gunicorn/app/wsgiapp.py", line 67, in run WSGIApplication("%(prog)s [OPTIONS] [APP_MODULE]", prog=prog).run() File "/usr/local/lib/python3.11/site-packages/gunicorn/app/base.py", line 236, in run super().run() File "/usr/local/lib/python3.11/site-packages/gunicorn/app/base.py", line 72, in run Arbiter(self).run() File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 229, in run self.halt(reason=inst.reason, exit_status=inst.exit_status) File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 342, in halt self.stop() File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 396, in stop time.sleep(0.1) File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 242, in handle_chld self.reap_workers() File "/usr/local/lib/python3.11/site-packages/gunicorn/arbiter.py", line 530, in reap_workers raise HaltServer(reason, self.WORKER_BOOT_ERROR) gunicorn.errors.HaltServer: <HaltServer 'Worker failed to boot.' 3>
我的配置文件
Dockerfile
FROM python:3.11.9-slim ENV PYTHONDONTWRITEBYTECODE 1 ENV PYTHONUNBUFFERED 1 WORKDIR /app RUN apt-get update \ && apt-get install -y --no-install-recommends \ gcc \ python3-dev \ musl-dev \ postgresql-client \ libpq-dev \ nginx \ supervisor \ && rm -rf /var/lib/apt/lists/* RUN rm /etc/nginx/nginx.conf COPY nginx.conf /etc/nginx/nginx.conf COPY requirements.txt /app/ RUN pip install --no-cache-dir -r requirements.txt COPY . /app/ RUN python manage.py collectstatic --no-input RUN python manage.py makemigrations RUN python manage.py migrate COPY supervisord.conf /etc/supervisor/conf.d/supervisord.conf EXPOSE 80 CMD ["/usr/bin/supervisord", "-c", "/etc/supervisor/conf.d/supervisord.conf"]
docker-compose.yml
services: web: build: . command: /usr/bin/supervisord -c /etc/supervisor/conf.d/supervisord.conf environment: DATABASE_NAME: ### DATABASE_USER: ### DATABASE_PASSWORD: ### DATABASE_HOST: ### DATABASE_PORT: 5432 volumes: - /path_to_static:/app/static #edited for stackoverflow - /path_to_static:/app/media #edited for stackoverflow networks: - backend nginx: build: context: . dockerfile: Dockerfile.nginx ports: - "80:80" volumes: - /path_to_static:/app/static #edited for stackoverflow - /path_to_static:/app/media #edited for stackoverflow networks: - backend depends_on: - web networks: backend:
Dockerfile.nginx
FROM nginx:latest RUN rm /etc/nginx/nginx.conf COPY nginx.conf /etc/nginx/nginx.conf
解决步骤
1. 确认Supervisor的directory字段配置
你的directory=/app是正确的——项目已复制到容器的/app目录,且Dockerfile中设置了WORKDIR /app,这部分无需修改。
2. 定位Gunicorn启动失败的具体原因
Gunicorn退出码3表示Worker启动失败,但当前日志未展示Django层面的报错。修改Supervisor中gunicorn的命令,开启调试日志:
command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000 --log-level debug
重新启动后,查看gunicorn.err.log就能看到具体错误(比如Django配置错误、数据库连接失败等)。
3. 修复Dockerfile中的核心问题
- 迁移时机错误:Docker build阶段无法访问docker-compose网络中的数据库,因此不能在build时执行
makemigrations和migrate。移除Dockerfile中的这两行:
改用entrypoint脚本在容器启动时执行迁移:RUN python manage.py makemigrations RUN python manage.py migrate
创建entrypoint.sh文件:
在Dockerfile中添加:#!/bin/sh python manage.py makemigrations python manage.py migrate python manage.py collectstatic --no-input exec "$@"COPY entrypoint.sh /app/ RUN chmod +x /app/entrypoint.sh ENTRYPOINT ["/app/entrypoint.sh"] - 重复部署Nginx:web容器和docker-compose中各部署了一个Nginx,会导致资源冲突。保留独立的Nginx服务,移除web容器中的Nginx相关配置:
修改Dockerfile的安装命令,去掉nginx:
同时删除Supervisor配置中的RUN apt-get update \ && apt-get install -y --no-install-recommends \ gcc \ python3-dev \ musl-dev \ postgresql-client \ libpq-dev \ supervisor \ && rm -rf /var/lib/apt/lists/*[program:nginx]块。
4. 验证数据库连接
确认docker-compose中的数据库环境变量正确,且数据库服务在同一个backend网络中。可以在容器启动后手动测试连接:
docker-compose exec web psql -h $DATABASE_HOST -U $DATABASE_USER -d $DATABASE_NAME
5. 优化Supervisor配置
- 提前创建日志目录,避免启动报错:
在Dockerfile中添加:RUN mkdir -p /var/log/gunicorn - 为Gunicorn添加稳定性参数:
command=gunicorn trade_calls.wsgi:application --bind 0.0.0.0:8000 --workers 2 --timeout 30 --log-level debug
内容的提问来源于stack exchange,提问作者Uday Singh Khalsa
相关产品推荐
相关产品推荐

