SSH远程训练神经网络遇Broken pipe错误,求无sudo权限解决方案
解决SSH断连导致训练中断的无sudo方案
以下是无需服务器sudo权限的解决办法,按可靠性排序:
1. 使用终端复用工具(最推荐)
这类工具能让训练进程在服务器后台独立运行,完全不受SSH连接断开的影响,重新连接后还能恢复会话查看输出。
screen:
- 连接服务器后,输入
screen创建新会话 - 在会话里运行训练脚本(比如
python train.py) - 需要暂时断开时,按
Ctrl+A+Ddetach会话 - 重新连接服务器后,输入
screen -r恢复会话;如果有多个会话,用screen -ls查看会话ID,再用screen -r <会话ID>恢复
- 连接服务器后,输入
tmux(功能比screen更丰富):
- 连接服务器后,输入
tmux new -s train_session创建命名会话 - 运行训练脚本
- 按
Ctrl+B+Ddetach会话 - 重新连接后,输入
tmux attach -t train_session恢复会话;用tmux ls查看所有会话
- 连接服务器后,输入
nohup(轻量后台运行):
- 直接在终端输入:
nohup python train.py > train.log 2>&1 &nohup让进程忽略挂断信号> train.log 2>&1把标准输出和错误输出都写入train.log文件&让进程后台运行
- 查看实时输出:
tail -f train.log - 重新连接后,用
ps aux | grep train.py检查进程是否在运行,继续用tail -f train.log查看日志
- 直接在终端输入:
2. 优化本地SSH客户端的心跳配置
如果之前改的是服务器端配置,试试修改本地的SSH参数,强制客户端定期发送心跳包,避免被中间网络设备判定为空闲连接而断开。
Linux/macOS本地终端:
编辑本地~/.ssh/config文件(没有就创建),添加:Host your_server_alias HostName xxx.xx.x.xxx User your_username ServerAliveInterval 30 ServerAliveCountMax 10 IPQoS throughputServerAliveInterval 30:每30秒向服务器发送一次心跳ServerAliveCountMax 10:连续10次心跳失败才断开连接IPQoS throughput:优先保障吞吐量,避免心跳包被限流
PuTTY:
- 打开PuTTY,在左侧导航栏点击
Connection - 勾选
Enable TCP keepalives (SO_KEEPALIVE option) - 设置
Seconds between keepalives为30 - 点击
Session,保存当前配置,下次连接直接用这个配置
- 打开PuTTY,在左侧导航栏点击
Visual Studio Code:
打开VS Code的SSH配置文件(快捷键Ctrl+Shift+P,输入Remote-SSH: Open Configuration File),添加和本地终端相同的参数:Host xxx.xx.x.xxx HostName xxx.xx.x.xxx User your_username ServerAliveInterval 30 ServerAliveCountMax 10
3. 直接带参数启动SSH连接
如果不想修改配置文件,每次连接时直接用ssh命令带参数:
ssh -o ServerAliveInterval=30 -o ServerAliveCountMax=10 -o IPQoS=throughput user@xxx.xx.x.xxx
内容的提问来源于stack exchange,提问作者NoProg
相关产品推荐
相关产品推荐

