Docker restart:always策略在Pumba杀容器后未自动重启问题
问题:Docker Compose配置restart:always但容器被Pumba杀死后未重启
我已为Docker Compose中的所有服务配置了restart:"always"重启策略,预期混沌工程工具Pumba随机杀死某个容器时,该容器会自动重启,但实际容器退出后并未重启。
相关信息
Pumba执行日志
./pumba --interval=1m --random -l info kill --signal=SIGTERM "re2:^example-voting-app_vote" INFO[0000] killing container dryrun=false id=5fe93ff11e5a6fbb9dc584848159341bc1fb710c747599fa03276f234a3faa19 name=/example-voting-app_vote_1 signal=SIGTERM
容器日志
vote_1 | WARNING: This is a development server. Do not use it in a production deployment. Use a production WSGI server instead. vote_1 | * Running on all addresses (0.0.0.0) vote_1 | * Running on http://127.0.0.1:80 vote_1 | * Running on http://172.20.0.5:80 vote_1 | Press CTRL+C to quit vote_1 | * Restarting with stat worker_1 | Connected to db worker_1 | Connecting to redis worker_1 | Found redis at 172.20.0.2 result_1 | [nodemon] 2.0.22 result_1 | [nodemon] to restart at any time, enter `rs` result_1 | [nodemon] watching path(s): *.* result_1 | [nodemon] watching extensions: js,mjs,json result_1 | [nodemon] starting `node server.js` vote_1 | * Debugger is active! vote_1 | * Debugger PIN: 861-258-488 result_1 | Sun, 02 Jul 2023 00:48:49 GMT body-parser deprecated bodyParser: use individual json/urlencoded middlewares at server.js:73:9 result_1 | Sun, 02 Jul 2023 00:48:49 GMT body-parser deprecated undefined extended: provide extended option at ../node_modules/body-parser/index.js:104:29 result_1 | App running on port 80 result_1 | Connected to db example-voting-app_vote_1 exited with code 0
Docker Compose配置文件
# version is now using "compose spec" # v2 and v3 are now combined! # docker-compose v1.27+ required services: vote: build: ./vote # use python rather than gunicorn for local dev command: python app.py depends_on: redis: condition: service_healthy healthcheck: test: ["CMD", "curl", "-f", "http://localhost"] interval: 15s timeout: 5s retries: 3 start_period: 10s volumes: - ./vote:/app ports: - "5000:80" networks: - front-tier - back-tier restart: "always" result: build: ./result # use nodemon rather than node for local dev entrypoint: nodemon server.js depends_on: db: condition: service_healthy volumes: - ./result:/app ports: - "5001:80" - "5858:5858" networks: - front-tier - back-tier restart: "always" worker: build: context: ./worker depends_on: redis: condition: service_healthy db: condition: service_healthy networks: - back-tier restart: "always" redis: image: redis:alpine volumes: - "./healthchecks:/healthchecks" healthcheck: test: /healthchecks/redis.sh interval: "5s" networks: - back-tier restart: "always" db: image: postgres:15-alpine environment: POSTGRES_USER: "postgres" POSTGRES_PASSWORD: "postgres" volumes: - "db-data:/var/lib/postgresql/data" - "./healthchecks:/healthchecks" healthcheck: test: /healthchecks/postgres.sh interval: "5s" networks: - back-tier restart: "always" # this service runs once to seed the database with votes # it won't run unless you specify the "seed" profile # docker compose --profile seed up -d seed: build: ./seed-data profiles: ["seed"] depends_on: vote: condition: service_healthy networks: - front-tier restart: "no" volumes: db-data: networks: front-tier: back-tier:
分析与解决方案
从容器日志可以看到,vote_1容器最终以退出码0正常退出。虽然restart: always理论上会重启任何状态退出的容器,但实际未触发重启可能由以下原因导致,对应解决方案如下:
1. 验证重启策略是否生效
首先确认容器的重启配置是否正确应用:
docker inspect example-voting-app_vote_1 | grep -A3 RestartPolicy
如果输出中Name不是always,说明配置未生效,需重新部署服务:
docker compose up -d --force-recreate vote
2. 测试手动停止容器的行为
执行手动停止命令,观察容器是否自动重启:
docker stop example-voting-app_vote_1
- 如果容器重启,说明Pumba发送的
SIGTERM信号导致应用进程优雅退出后,Docker的重启逻辑出现异常; - 如果容器未重启,说明Docker daemon的重启策略未正常工作,需检查Docker服务状态或配置。
3. 更换Pumba发送的信号类型
Flask开发服务器收到SIGTERM会优雅退出并返回0,尝试用SIGKILL强制杀死进程(退出码非0),验证是否触发重启:
./pumba --interval=1m --random -l info kill --signal=SIGKILL "re2:^example-voting-app_vote"
4. 查看Docker daemon日志排查问题
检查Docker守护进程的日志,寻找容器重启失败或未触发的原因:
# 系统d环境 journalctl -u docker.service -f # 非系统d环境 tail -f /var/log/docker.log
5. 更换应用启动方式
开发服务器的信号处理逻辑可能与生产环境不同,将vote服务的启动命令改为生产级WSGI服务器(如gunicorn)测试:
修改Compose中vote服务的command:
command: gunicorn -w 4 -b 0.0.0.0:80 app:app
重新部署后再用Pumba测试。
内容的提问来源于stack exchange,提问作者akanesora
相关产品推荐
相关产品推荐

