You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Ubuntu 22.04.1树莓派Docker Swarm无法部署Stack及创建Overlay网络

Docker Swarm栈部署异常排查与解决

问题描述

我有一个由7台树莓派组成的小型Docker Swarm集群:

ubuntu@rpi105:~/stacks$ docker node ls
ID                            HOSTNAME   STATUS    AVAILABILITY   MANAGER STATUS   ENGINE VERSION
me5ma5mtkl98iutcztnyar7nf     rpi102     Ready     Active                          20.10.23
125x7tjps3om8qp4awmlt9nkn     rpi103     Ready     Active                          20.10.23
cxoluolounydb8wxydfhd0pd3     rpi104     Ready     Active                          20.10.23
6psveckpp209kx29je9bdug67 *   rpi105     Ready     Active         Leader           20.10.23
tva9hlxlsgsagoic92b5nsv9e     rpi106     Ready     Active         Reachable        20.10.23
qu2wboooaoux1kiy86yw2nkdk     rpi107     Ready     Active                          20.10.23
iu3nnacxqlz34lgy2tzzczxf2     rpi108     Ready     Active         Reachable        20.10.23

命令行部署服务正常,但使用Stack文件部署时会卡住。测试用基础Stack文件如下:

version: "3.9"

services:
  nginx_test:
    image: nginx:latest
    deploy:
      replicas: 1
    ports:
      - 81:80

部署后服务长期处于“New”状态,无报错:

ubuntu@rpi105:~/stacks$ docker stack deploy -c test_nginx.yml testng
Creating network testng_default
Creating service testng_nginx_test
ubuntu@rpi105:~/stacks$ docker stack ps testng
ID             NAME                  IMAGE          NODE      DESIRED STATE   CURRENT STATE        ERROR     PORTS
owgslvqtbgfx   testng_nginx_test.1   nginx:latest             Running         New 21 minutes ago

自动创建的网络无驱动:

ubuntu@rpi105:~/stacks$ docker network ls
NETWORK ID     NAME             DRIVER    SCOPE
b19cd2106cf0   bridge           bridge    local
d6ecdf2de829   host             host      local
nys70xbvgset   ingress          overlay   swarm
46fa0761429f   none             null      local
mhggl0kyq5o5   testng_default             swarm

手动创建Overlay网络同样无驱动:

ubuntu@rpi105:~/stacks$ docker network create -d overlay testnet
j5pg96332w5hvy7qoratpyvzc
ubuntu@rpi105:~/stacks$ docker network ls
NETWORK ID     NAME             DRIVER    SCOPE
b19cd2106cf0   bridge           bridge    local
d6ecdf2de829   host             host      local
nys70xbvgset   ingress          overlay   swarm
46fa0761429f   none             null      local
j5pg96332w5h   testnet                    swarm
mhggl0kyq5o5   testng_default             swarm
ubuntu@rpi105:~/stacks$ docker network inspect testnet
[
    {
        "Name": "testnet",
        "Id": "j5pg96332w5hvy7qoratpyvzc",
        "Created": "2023-02-02T19:30:36.372991984Z",
        "Scope": "swarm",
        "Driver": "",
        "EnableIPv6": false,
        "IPAM": {
            "Driver": "",
            "Options": null,
            "Config": null
        },
        "Internal": false,
        "Attachable": false,
        "Ingress": false,
        "ConfigFrom": {
            "Network": ""
        },
        "ConfigOnly": false,
        "Containers": null,
        "Options": null,
        "Labels": null
    }
]

集群所有节点环境:

  • Ubuntu 22.04.1 LTS
  • Docker Engine - Community 20.10.23
  • Docker Compose version v2.15.1

解决方案

1. 检查Docker服务状态

所有节点上执行,确保Docker服务正常运行:

sudo systemctl status docker

若有异常,重启服务:

sudo systemctl restart docker

Worker节点重启后需重新加入Swarm集群,Manager节点重启后确认集群状态。

2. 打通集群节点必备端口

Overlay网络依赖以下端口通信,确保所有节点间这些端口开放:

  • TCP 2377(集群管理)
  • TCP/UDP 7946(节点 gossip 通信)
  • UDP 4789(Overlay数据传输)

用nc测试端口连通性(示例:从rpi102测试rpi105的2377端口):

nc -zv rpi105 2377

若端口不通,调整UFW防火墙规则:

sudo ufw allow 2377/tcp
sudo ufw allow 7946/tcp
sudo ufw allow 7946/udp
sudo ufw allow 4789/udp
sudo ufw reload

3. 重新指定Stack使用Ingress网络

修改Stack文件,明确使用已有的ingress网络,避免自动创建异常网络:

version: "3.9"

networks:
  default:
    external:
      name: ingress

services:
  nginx_test:
    image: nginx:latest
    deploy:
      replicas: 1
    ports:
      - 81:80

4. 检查并切换存储驱动

树莓派上需使用overlay2存储驱动,检查当前驱动:

docker info | grep "Storage Driver"

若不是,创建/修改/etc/docker/daemon.json:

{
  "storage-driver": "overlay2"
}

重启Docker服务:

sudo systemctl restart docker

5. 重置Swarm集群(最后手段)

若以上操作无效,备份服务配置后,在Leader节点重置Swarm:

docker swarm leave --force
docker swarm init --advertise-addr <Leader节点IP>

随后让所有Worker和其他Manager节点重新加入集群。

内容的提问来源于stack exchange,提问作者juan noguera

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 11:46:07