Docker容器执行network disconnect后仍能通信的问题排查求助
Let me break down why your network disconnect simulation didn't work as expected, and walk you through practical steps to diagnose the issue:
Possible Causes
1. Established TCP Connections Don't Drop Immediately
TCP is a connection-oriented protocol. Once node1 and node2 have an active TCP connection set up, removing node2 from the dm_default network won't instantly tear that connection down. The link will stay alive until one side tries to send data and hits a failure (like a timeout or ICMP unreachable message), or until TCP keepalive mechanisms detect the break. If your apps are idle after the network is disconnected, they might not realize the link is down yet.
2. Your Apps Are Using Host Port Mapping Instead of Container Network
Looking at your docker-compose.yaml, node1 has a port mapping - "5500:5500". If node2 is configured to connect to your host machine's IP (e.g., your local network IP or localhost) on port 5500 instead of node1's container IP/service name, it’s bypassing the dm_default network entirely. This connection would route through your host’s network stack, so removing node2 from the custom network won’t affect this communication.
3. node2 Is Still Connected to Another Shared Network
Docker containers can attach to multiple networks at once. It’s possible node2 is part of another network (like the default bridge network) that node1 also belongs to. Removing it from dm_default would only cut off one path, not all possible routes between the two containers.
Troubleshooting Steps
- Verify
node2's network attachments: Rundocker inspect node2and check theNetworkSettings > Networkssection. Confirm there are no other networks listed thatnode1is also connected to. - Check the connection target in
node2: Exec intonode2withdocker exec -it node2 bash, then runss -tulpnto list active connections. Look at the destination IP of the port 5500 connection—if it’s your host’s IP instead of172.20.0.2, that’s the root cause. You can also runtraceroute 172.20.0.2to see if the path is blocked (or if it’s even using that IP). - Test active communication: Try sending a test payload from one app to the other after disconnecting the network. Idle connections won’t trigger TCP failure detection, but active data transfer will reveal if the link is truly still functional.
- Inspect traffic with
tcpdump: Runtcpdump -i any port 5500inside both containers to check if packets are still being exchanged. If no new packets appear after the network disconnect, the connection is likely just in an idle "zombie" state waiting for a failure trigger.
内容的提问来源于stack exchange,提问作者Chirrut Imwe

