求解WongKinYiu/PyTorch_YOLOv4的RuntimeError设备不匹配问题
问题描述
使用WongKinYiu的PyTorch_YOLOv4仓库训练模型时触发设备不匹配错误,YOLOv7社区已有相关解决方案,但YOLOv4相关内容缺失,特此求助。
报错栈信息
Traceback (most recent call last):
File "/content/PyTorch_YOLOv4/train.py", line 537, in
train(hyp, opt, device, tb_writer, wandb)
File "/content/PyTorch_YOLOv4/train.py", line 288, in train
loss, loss_items = compute_loss(pred, targets.to(device), model) # loss scaled by batch_size
File "/content/PyTorch_YOLOv4/utils/loss.py", line 69, in compute_loss
tcls, tbox, indices, anchors = build_targets(p, targets, model) # targets
File "/content/PyTorch_YOLOv4/utils/loss.py", line 151, in build_targets
a, t = at[j], t.repeat(na, 1, 1)[j] # filter
RuntimeError: indices should be either on cpu or on the same device as the indexed tensor (cpu)
报错核心是索引张量j和被索引的t.repeat(na,1,1)设备不一致(一个在GPU、一个在CPU),只需统一两者设备即可。
修改utils/loss.py中build_targets函数的对应代码:
找到第151行附近的代码,将重复后的t张量移动到j所在的设备,替换原代码为:
t_repeated = t.repeat(na, 1, 1).to(j.device) a, t = at[j], t_repeated[j] # filter
或者更彻底的方式:在build_targets函数开头获取目标张量的设备(比如device = targets.device),后续生成的所有索引类张量(如j、at等)都显式指定该设备,确保全程设备一致。
内容的提问来源于stack exchange,提问作者MheadHero

