Docker Swarm部署容器频繁重启问题排查求助
容器在Docker Swarm中反复重启的排查问题
环境与配置
通过docker node update初始化并更新Docker Swarm网络后,部署了如下容器:
Docker Compose 文件
version: '3.8' networks: test: driver: overlay attachable: true services: identity_vdr: networks: - test image: identity_vdr:1.0 ports: - "9701:9701" - "9702:9702" deploy: replicas: 1 placement: constraints: [node.hostname == indy1]
Dockerfile 内容
# Uses the base image of Ubuntu 16.04 FROM ubuntu:16.04 # Sets the non-interactive environment to avoid issues with prompts ENV DEBIAN_FRONTEND=noninteractive # Updates the system and installs dependencies RUN apt-get update -y && \ apt-get install -y \ git \ wget \ vim \ python3.5 \ python3-pip \ python-setuptools \ python3-nacl \ apt-transport-https \ ca-certificates # Updates pip and setuptools RUN pip3 install -U 'pip<10.0.0' setuptools==44.0.0 # Adds the necessary keys for the Sovrin repositories RUN apt-key adv --keyserver keyserver.ubuntu.com --recv-keys CE7709D068DB5E88 && \ apt-key adv --keyserver keyserver.ubuntu.com --recv-keys BD33704C # Adds the Sovrin repositories RUN echo "deb https://repo.sovrin.org/deb xenial master" >> /etc/apt/sources.list && \ echo "deb https://repo.sovrin.org/sdk/deb xenial master" >> /etc/apt/sources.list # Updates the system again and installs the Hyperledger Indy packages RUN apt-get update -y && \ apt-get install -y \ indy-node=1.13.0~dev1213 \ libindy-crypto=0.4.5 \ python3-indy-crypto=0.4.5 \ python3-orderedset=2.0 \ python3-psutil=5.4.3 \ python3-pympler=0.5 \ indy-plenum=1.13.0~dev1021 \ libindy=1.15.0~1536-xenial \ indy-cli=1.15.0~1536-xenial # Installs the python3-indy package RUN pip3 install python3-indy==1.15.0 # Clones the GitHub repository # RUN git clone https://github.com/timo-kang/Hyperledger-Indy-Tutorial.git /opt/Hyperledger-Indy-Tutorial COPY . /opt/Hyperledger-Indy-Tutorial # Updates the Indy configuration file RUN awk '{if (index($1, "NETWORK_NAME") != 0) {print("NETWORK_NAME = \"sandbox\"")} else print($0)}' /etc/indy/indy_config.py > /tmp/indy_config.py && \ mv /tmp/indy_config.py /etc/indy/indy_config.py # Sets the working directory WORKDIR /opt/Hyperledger-Indy-Tutorial EXPOSE 9701 9702 # Sets the default command CMD ["bash"]
问题现象
本地构建并运行容器时一切正常,但部署到Swarm栈后,容器每隔几秒就停止并自动重建,无错误提示:
ID NAME IMAGE NODE DESIRED STATE CURRENT STATE ERROR PORTS g8exkoszfivi node-1-stack_identity_vdr.1 identity_vdr:1.0 indy1 Running Starting 1 second ago o7an7o4r4cnj \_ node-1-stack_identity_vdr.1 identity_vdr:1.0 indy1 Shutdown Complete 6 seconds ago ogwtjvct1d26 \_ node-1-stack_identity_vdr.1 identity_vdr:1.0 indy1 Shutdown Complete 13 seconds ago qu0is4xgl3zq \_ node-1-stack_identity_vdr.1 identity_vdr:1.0 indy1 Shutdown Complete 21 seconds ago b7z9k3aqxok1 \_ node-1-stack_identity_vdr.1 identity_vdr:1.0 indy1 Shutdown Complete 27 seconds ago
曾怀疑Dockerfile中的CMD ["bash"]是问题所在,但移除后情况并未改善。且在Swarm栈中无法通过docker logs <ID>查看日志,需要有效的排查方法。
排查方法
1. 查看Swarm服务日志
直接通过服务名称获取所有实例(包括已停止的)的日志,命令:
docker service logs -f node-1-stack_identity_vdr
-f参数可以实时跟踪日志输出,能看到容器启动、退出的详细信息。
2. 检查容器退出码
通过容器ID查看退出状态码,判断退出原因:
docker inspect <container-id> | grep -A5 "State"
重点关注ExitCode字段:
0:容器正常退出(通常是启动命令执行完毕)1:一般应用错误137:容器被SIGKILL信号终止(大概率是内存不足触发OOM Killer)143:容器被SIGTERM信号终止(正常停止信号)
3. 临时修改启动命令,进入容器调试
在Compose文件的services.identity_vdr下添加command字段,让容器保持运行:
services: identity_vdr: # ... 其他配置 ... command: ["sleep", "infinity"]
重新部署后,进入容器手动检查环境:
docker exec -it <container-id> bash
在容器内可以:
- 手动启动目标服务,查看报错信息
- 验证配置文件(如
/etc/indy/indy_config.py)是否正确 - 检查依赖包是否正常安装
- 测试端口
9701/9702是否能正常监听
4. 检查Swarm节点资源
登录到indy1节点,检查CPU、内存、磁盘资源是否充足:
# 查看节点状态 docker node inspect indy1 # 查看资源占用 top # 检查是否有OOM Killer记录 dmesg | grep -i oom
如果节点内存不足,容器可能会被系统强制终止。
5. 验证Overlay网络状态
检查Swarm overlay网络是否正常,容器是否能正常接入:
docker network inspect node-1-stack_test
查看Containers字段是否包含目标容器,Peers字段是否包含所有Swarm节点。
6. 跟踪Swarm服务事件
实时查看服务的状态变化事件,了解重启触发原因:
docker service events node-1-stack_identity_vdr
内容的提问来源于stack exchange,提问作者brenodacosta
相关产品推荐
相关产品推荐

