自定义数据集训练YOLOv7遇torch.distributed导入错误求解决方案
解决YOLOv7训练时torch.distributed导入失败问题
问题场景
执行以下命令训练YOLOv7模型:
python train.py --img-size 2048 --cfg cfg/training/yolov7.yaml --hyp data/road_sign_data.yaml --batch 8 --epochs 100 --data data/road_sign.yaml --weights yolov7_training.pt --workers 24 --name yolo_road_det
出现错误:
Traceback (most recent call last): File "C:\Users\531558\Documents\streamline2\yolov7\train.py", line 12, in import torch.distributed as dist File "C:\Users\531558\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\torch\__init__.py", line 141, in raise err OSError: [WinError 126] The specified module could not be found. Error loading "C:\Users\531558\AppData\Local\Packages\PythonSoftwareFoundation.Python.3.11_qbz5n2kfra8p0\LocalCache\local-packages\Python311\site-packages\torch\lib\shm.dll" or one of its dependencies.
已尝试将Python版本从3.12改为3.11.9,问题仍未解决,其他YOLOv7训练方法也无效。
解决方案
- 安装VC运行库:
shm.dll依赖Visual C Redistributable,下载并安装对应版本的Microsoft Visual C++ Redistributable Packages(推荐2019或2022版),安装后重启系统再尝试训练。 - 重新安装适配的PyTorch:卸载现有PyTorch组件,使用官方命令重新安装适配Windows+Python3.11的版本:
pip uninstall torch torchvision torchaudio -y # GPU版本(以CUDA12.1为例) pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu121 # CPU版本 pip install torch torchvision torchaudio - 禁用分布式训练(单设备场景):如果仅用单GPU或CPU训练,修改
train.py:
找到导入torch.distributed的代码段,改为:
同时运行命令时添加import torch try: import torch.distributed as dist except ImportError: dist = None--device cpu或--device 0指定单设备,部分YOLOv7版本会自动禁用分布式逻辑。 - 检查依赖完整性:用Dependency Walker工具查看
shm.dll的依赖项,补充缺失的系统DLL;若shm.dll本身缺失,重新安装PyTorch即可。 - 重建干净虚拟环境:创建新的Python虚拟环境,先安装YOLOv7的
requirements.txt依赖,再安装适配的PyTorch版本,避免依赖冲突。
内容的提问来源于stack exchange,提问作者Garance MARION
相关产品推荐
相关产品推荐

