You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

YOLO模型训练遇inplace操作引发RuntimeError,求排查方法

训练Yolov5时loss.backward()触发RuntimeError(inplace操作问题)

训练Yolov5模型时,执行loss.backward()触发RuntimeError,提示梯度计算所需的某一变量被inplace操作修改(版本从1变为2),以下是训练代码、报错信息及定位解决方法:

训练代码

def main():
    # Initialize model, loss, and optimizer
    model = Yolov5(version='l')
    print(f"{sum(p.numel() for p in model.parameters())/1e6} million parameters")
    criterion = YOLO_LOSS(model, rect_training=True)
    optimizer = AdamW(model.parameters(), lr=0.5)
    
    # Number of epochs
    num_epochs = 5
   
    # Load data (train_loader, val_loader)
    train_loader, val_loader = get_loaders(db_root_dir="./datasets", batch_size=2, num_classes=2)
    
    # Training loop
    for epoch in range(num_epochs):
        model.train()
        epoch_loss = 0
    
        for imgs, targets in train_loader:
            # device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
            # # Use the device (CPU or GPU)
            # imgs = imgs.float().to(device)  # Move to GPU if available
            # targets = targets.to(device)
    
            # Aktifkan deteksi anomali sebelum backward pass
            torch.autograd.set_detect_anomaly(True)
            # Forward pass
            outputs = model(imgs)
            loss = criterion(outputs, targets, pred_size=imgs.shape[2:4])  # cls_loss + box_loss + dfl_loss
    
            optimizer.zero_grad()
            loss.backward()
            optimizer.step()
    
            epoch_loss += loss.item()
    
        print(f"Epoch {epoch+1}/{num_epochs} | Loss: {epoch_loss/len(train_loader)}")

报错信息

Exception has occurred: RuntimeError
one of the variables needed for gradient computation has been modified by an inplace 
operation: [torch.FloatTensor [2, 3, 20, 20]] is at version 2; expected version 1 instead.
Hint: the backtrace further above shows the operation that failed to compute its gradient.
The variable in question was changed in there or anywhere later. Good luck!
  File "C:\Users\HANIF\1_TA\siap\train.py", line 39, in main
    loss.backward()
  File "C:\Users\HANIF\1_TA\siap\train.py", line 47, in <module>
    main()
RuntimeError: one of the variables needed for gradient computation has been modified by an 
inplace operation: [torch.FloatTensor [2, 3, 20, 20]] is at version 2; expected version 1
instead.
Hint: the backtrace further above shows the operation that failed to compute its gradient. 
The variable in question was changed in there or anywhere later. Good luck!

定位与解决方案

定位inplace操作

inplace操作指直接修改张量本身而非创建新副本的操作(如x += y、x[0]=1、torch.nn.ReLU(inplace=True)),可通过以下方式定位:

  • 你已开启torch.autograd.set_detect_anomaly(True),运行代码时会输出带Anomaly detected标记的详细回溯,直接指向引发问题的inplace操作位置。
  • 重点检查YOLO_LOSS类的实现:损失计算过程中是否直接修改了outputs输入张量,或使用了带inplace参数的操作。
  • 检查Yolov5模型的forward方法:确认模型内部是否存在inplace=True的层(如激活层、归一化层),或手动的张量修改操作。

解决方法

  1. 禁用inplace参数:
    将模型中所有带inplace=True的层(如ReLU、SiLU)改为inplace=False,避免直接修改梯度依赖的张量。
  2. 替换inplace赋值操作:
    将x += y改为x = x + y,将x[mask] = value改为x = x.clone(); x[mask] = value,通过创建副本保留原始张量的梯度信息。
  3. 检查损失函数实现:
    确保损失计算过程中不修改输入的outputs张量,所有变换操作基于副本进行,例如先执行outputs_clone = outputs.clone()再处理。
  4. 统一设备配置:
    取消注释代码中的设备转移逻辑,确保模型、数据、损失函数都运行在同一设备(CPU/GPU)上,避免跨设备操作引发隐性问题。
  5. 确认梯度流程顺序:
    保证optimizer.zero_grad()在loss.backward()之前调用(你的代码已满足,但需确认无其他提前修改梯度的操作)。

内容的提问来源于stack exchange,提问作者Hanif Al-Farisi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 01:44:53