PyTorch自定义物理损失函数反向传播遇原地操作错误
错误详情
Exception has occurred: RuntimeError
one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor []], which is output 0 of AsStridedBackward0, is at version 2832; expected version 2831 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck!
File "...\GNN_OPF\src\train.py", line 60, in train_batch
loss.backward()
File "...\GNN_OPF\src\train.py", line 33, in train_model
epoch_loss += train_batch(data=batch, model=gnn, optimizer=optimizer, criterion=criterion)
File "...\GNN_OPF\src\main.py", line 48, in main
model, losses, val_losses = train_model(arguments, train, val, test)
File "...\GNN_OPF\src\main.py", line 58, in
main()
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.FloatTensor []], which is output 0 of AsStridedBackward0, is at version 2832; expected version 2831 instead. Hint: the backtrace further above shows the operation that failed to compute its gradient. The variable in question was changed in there or anywhere later. Good luck!
错误触发场景
仅当physics_crit=True时触发,对应train_batch函数:
def train_batch(data, model, optimizer, criterion, physics_crit=True, device='cpu'): model.to(device) th.autograd.set_detect_anomaly(True) optimizer.zero_grad() out = model(data) if physics_crit: loss = physics_loss(data, out) else: loss = criterion(out, data.y) loss.backward() #Error pops up when this line is called
问题来源函数(physics_loss)
def physics_loss(network, output, log_loss=True): active_imbalance = output[:,0] # th.zeros(output_r.shape[0]) reactive_imbalance = output[:,1]#th.zeros(output_r.shape[0]) resist_line_total = th.multiply(network.edge_attr[:, 0], network.edge_attr[:, -1]) react_line_total = th.multiply(network.edge_attr[:, 1], network.edge_attr[:, -1]) denom = th.add(th.multiply(resist_line_total, resist_line_total),th.multiply(react_line_total, react_line_total)) conductances = th.div(resist_line_total , denom) susceptances = th.multiply(th.div(react_line_total, denom),-1.0) for i, x in enumerate(th.transpose(network.edge_index, 0, 1)): angle_diff = th.sub(output[x[0],3] , output[x[1],3]) active_imbalance[x[0]] =th.sub(active_imbalance[x[0]],th.multiply(th.multiply(th.abs_(output[x[0], 2]) , th.abs_(output[x[1], 2])), th.add(th.multiply(conductances[i], th.cos(angle_diff)) , th.multiply(susceptances[i], th.sin(angle_diff))))) reactive_imbalance[x[0]] = th.sub(reactive_imbalance[x[0]],th.multiply( th.multiply(th.abs_(output[x[0], 2]), th.abs_(output[x[1], 2])), th.sub(th.multiply(conductances[i], th.sin(angle_diff)) , th.multiply(susceptances[i], th.cos(angle_diff))))) if log_loss: tot_loss = th.log(th.add(1.0 , th.sum(th.add(th.multiply(active_imbalance, active_imbalance), th.multiply(reactive_imbalance, reactive_imbalance))))) else: tot_loss = th.sum(np.abs(active_imbalance) + np.abs(reactive_imbalance)) return tot_loss
已尝试操作:全程使用PyTorch函数替代原生运算,detach()测试时无错误。
问题根源
- 原地修改计算图张量:
th.abs_()是原地操作,直接修改了模型输出output的元素,破坏了梯度追踪的计算图。 - 视图张量的原地赋值:
active_imbalance = output[:,0]是output的切片视图,而非独立张量,后续对active_imbalance[x[0]]的赋值会间接修改output的原始数据。 - 跨框架操作脱离计算图:
log_loss=False分支使用np.abs将PyTorch张量转为NumPy数组,导致梯度无法传播。
修复方案
1. 替换所有原地操作
将带下划线的PyTorch原地函数(如th.abs_())改为非原地版本(th.abs()),避免修改原始张量。
2. 克隆初始张量
对output的切片进行克隆,创建独立张量,切断与原始output的视图关联:
active_imbalance = output[:,0].clone() reactive_imbalance = output[:,1].clone()
3. 统一使用PyTorch函数
将np.abs替换为th.abs,确保全程在PyTorch计算图内操作。
修复后的physics_loss示例
def physics_loss(network, output, log_loss=True): # 克隆输出切片,避免修改原始output active_imbalance = output[:,0].clone() reactive_imbalance = output[:,1].clone() resist_line_total = network.edge_attr[:, 0] * network.edge_attr[:, -1] react_line_total = network.edge_attr[:, 1] * network.edge_attr[:, -1] denom = resist_line_total ** 2 + react_line_total ** 2 conductances = resist_line_total / denom susceptances = - (react_line_total / denom) # 预先转置边索引,减少循环内重复操作 edge_index = network.edge_index.transpose(0, 1) for i, (u, v) in enumerate(edge_index): angle_diff = output[u, 3] - output[v, 3] # 使用非原地abs,不修改原始output u_abs = th.abs(output[u, 2]) v_abs = th.abs(output[v, 2]) term_active = u_abs * v_abs * (conductances[i] * th.cos(angle_diff) + susceptances[i] * th.sin(angle_diff)) term_reactive = u_abs * v_abs * (conductances[i] * th.sin(angle_diff) - susceptances[i] * th.cos(angle_diff)) # 用减法更新张量,避免原地赋值 active_imbalance[u] -= term_active reactive_imbalance[u] -= term_reactive if log_loss: tot_loss = th.log(1.0 + th.sum(active_imbalance ** 2 + reactive_imbalance ** 2)) else: # 替换numpy.abs为PyTorch版本 tot_loss = th.sum(th.abs(active_imbalance) + th.abs(reactive_imbalance)) return tot_loss
额外建议
- 保持
th.autograd.set_detect_anomaly(True)开启,它能在错误发生前定位具体的原地操作位置。 - 尽量避免在梯度计算流程中使用原地操作,PyTorch的非原地操作会自动维护计算图的完整性。
内容的提问来源于stack exchange,提问作者B0B

