MacBook M2 Air上YOLOv5基于MPS训练报错求助
问题场景
已配置PyTorch的MPS后端实现GPU加速,但在MacBook M2 Air上运行YOLOv5训练时触发报错,执行的训练命令如下:
RES_DIR = set_res_dir() if TRAIN: !python /Users/krishpatel/yolov5/train.py --data /Users/krishpatel/yolov5/roboflow/data.yaml --weights yolov5s.pt \ --img 640 --epochs {EPOCHS} --batch-size 32 --device mps --name {RES_DIR}
报错信息:
UserWarning: The operator 'aten::nonzero' is not currently supported on the MPS backend and will fall back to run on the CPU. This may have performance implications. (Triggered internally at /Users/runner/work/pytorch/pytorch/pytorch/aten/src/ATen/mps/MPSFallback.mm:11.)
t = t[j] # filter
0%| | 0/20 [00:16<?, ?it/s]
Traceback (most recent call last):
File "/Users/krishpatel/yolov5/train.py", line 630, in
main(opt)
File "/Users/krishpatel/yolov5/train.py", line 524, in main
train(opt.hyp, opt, device, callbacks)
File "/Users/krishpatel/yolov5/train.py", line 307, in train
loss, loss_items = compute_loss(pred, targets.to(device)) # loss scaled by batch_size
File "/Users/krishpatel/yolov5/utils/loss.py", line 125, in call
tcls, tbox, indices, anchors = self.build_targets(p, targets) # targets
File "/Users/krishpatel/yolov5/utils/loss.py", line 213, in build_targets
j, k = ((gxy % 1 < g) & (gxy > 1)).T
NotImplementedError: The operator 'aten::remainder.Tensor_out' is not currently implemented for the MPS device. As a temporary fix, you can set the environment variablePYTORCH_ENABLE_MPS_FALLBACK=1to use the CPU as a fallback for this op. WARNING: this will be slower than running natively on MPS.
纯CPU训练速度极慢,希望找到无需改用Colab的可行方案。
解决建议
启用MPS fallback临时兼容
设置环境变量让不支持的算子自动回退到CPU执行,虽不如纯MPS快,但远胜纯CPU。有两种方式:- 命令行前缀设置:
PYTORCH_ENABLE_MPS_FALLBACK=1 python /Users/krishpatel/yolov5/train.py --data /Users/krishpatel/yolov5/roboflow/data.yaml --weights yolov5s.pt --img 640 --epochs {EPOCHS} --batch-size 32 --device mps --name {RES_DIR} - Python代码开头设置:
import os os.environ['PYTORCH_ENABLE_MPS_FALLBACK'] = '1'
- 命令行前缀设置:
修改YOLOv5代码替换不支持的算子
报错源于gxy % 1调用了MPS不支持的remainder算子,可替换为MPS支持的torch.frac()(功能等价,计算张量的小数部分)。打开/Users/krishpatel/yolov5/utils/loss.py,找到第213行:j, k = ((gxy % 1 < g) & (gxy > 1)).T修改为:
j, k = ((torch.frac(gxy) < g) & (gxy > 1)).T保存后重新运行训练,即可避免该报错。
升级PyTorch到最新版本
PyTorch对MPS的算子支持持续更新,新版本可能已适配aten::remainder.Tensor_out。执行以下命令升级:pip install --upgrade torch torchvision torchaudio调整训练参数优化性能
- 适当降低batch-size(如从32改为16或8),减少MPS内存占用,避免潜在的内存溢出问题;
- 改用更小的预训练模型(如
yolov5n.pt),计算量更小,训练速度更快; - 降低输入图像尺寸(如从640改为416),减少单步计算量,提升训练效率。
内容的提问来源于stack exchange,提问作者Krish Patel

