YOLO模型训练遇inplace操作引发RuntimeError,求排查方法
训练Yolov5时loss.backward()触发RuntimeError(inplace操作问题)
训练Yolov5模型时,执行loss.backward()触发RuntimeError,提示梯度计算所需的某一变量被inplace操作修改(版本从1变为2),以下是训练代码、报错信息及定位解决方法:
训练代码
def main(): # Initialize model, loss, and optimizer model = Yolov5(version='l') print(f"{sum(p.numel() for p in model.parameters())/1e6} million parameters") criterion = YOLO_LOSS(model, rect_training=True) optimizer = AdamW(model.parameters(), lr=0.5) # Number of epochs num_epochs = 5 # Load data (train_loader, val_loader) train_loader, val_loader = get_loaders(db_root_dir="./datasets", batch_size=2, num_classes=2) # Training loop for epoch in range(num_epochs): model.train() epoch_loss = 0 for imgs, targets in train_loader: # device = torch.device("cuda" if torch.cuda.is_available() else "cpu") # # Use the device (CPU or GPU) # imgs = imgs.float().to(device) # Move to GPU if available # targets = targets.to(device) # Aktifkan deteksi anomali sebelum backward pass torch.autograd.set_detect_anomaly(True) # Forward pass outputs = model(imgs) loss = criterion(outputs, targets, pred_size=imgs.shape[2:4]) # cls_loss + box_loss + dfl_loss optimizer.zero_grad() loss.backward() optimizer.step() epoch_loss += loss.item() print(f"Epoch {epoch+1}/{num_epochs} | Loss: {epoch_loss/len(train_loader)}")
报错信息
Exception has occurred: RuntimeError one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [2, 3, 20, 20]] is at version 2; expected version 1 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck! File "C:\Users\HANIF\1_TA\siap\train.py", line 39, in main loss.backward() File "C:\Users\HANIF\1_TA\siap\train.py", line 47, in <module> main() RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor [2, 3, 20, 20]] is at version 2; expected version 1 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck!
定位与解决方案
定位inplace操作
inplace操作指直接修改张量本身而非创建新副本的操作(如x += y、x[0]=1、torch.nn.ReLU(inplace=True)),可通过以下方式定位:
- 你已开启
torch.autograd.set_detect_anomaly(True),运行代码时会输出带Anomaly detected标记的详细回溯,直接指向引发问题的inplace操作位置。 - 重点检查
YOLO_LOSS类的实现:损失计算过程中是否直接修改了outputs输入张量,或使用了带inplace参数的操作。 - 检查Yolov5模型的forward方法:确认模型内部是否存在
inplace=True的层(如激活层、归一化层),或手动的张量修改操作。
解决方法
- 禁用inplace参数:
将模型中所有带inplace=True的层(如ReLU、SiLU)改为inplace=False,避免直接修改梯度依赖的张量。 - 替换inplace赋值操作:
将x += y改为x = x + y,将x[mask] = value改为x = x.clone(); x[mask] = value,通过创建副本保留原始张量的梯度信息。 - 检查损失函数实现:
确保损失计算过程中不修改输入的outputs张量,所有变换操作基于副本进行,例如先执行outputs_clone = outputs.clone()再处理。 - 统一设备配置:
取消注释代码中的设备转移逻辑,确保模型、数据、损失函数都运行在同一设备(CPU/GPU)上,避免跨设备操作引发隐性问题。 - 确认梯度流程顺序:
保证optimizer.zero_grad()在loss.backward()之前调用(你的代码已满足,但需确认无其他提前修改梯度的操作)。
内容的提问来源于stack exchange,提问作者Hanif Al-Farisi
相关产品推荐
相关产品推荐

