运行MolScribe训练脚本时遇反向传播原地操作RuntimeError
问题
运行MolScribe仓库中的bash scripts/train_uspto_joint_chartok_1m680k.sh脚本时触发RuntimeError,错误信息如下:
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.cuda.LongTensor [1, 123]] is at version 1; expected version 0 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).
错误发生在train.py的train_fn函数中scaler.scale(loss).backward()行,完整回溯信息:
Traceback (most recent call last): File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 612, in <module> main() File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 598, in main train_loop(args, train_df, valid_df, aux_df, tokenizer, args.save_path) File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 380, in train_loop avg_loss, global_step = train_fn( File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 221, in train_fn scaler.scale(loss).backward() File "/home/qk/anaconda3/envs/MolScribe/lib/python3.9/site-packages/torch/_tensor.py", line 488, in backward torch.autograd.backward( File "/home/qk/anaconda3/envs/MolScribe/lib/python3.9/site-packages/torch/autograd/__init__.py", line 197, in backward Variable._execution_engine.run_backward( # Calls into the C++ engine to run the backward pass RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.cuda.LongTensor [1, 123]] is at version 1; expected version 0 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).
已尝试添加with torch.autograd.set_detect_anomaly(True)调试,但回溯信息无变化,无法定位具体操作。
解决建议
- 正确启用异常检测:不要仅在
backward()附近添加检测,把torch.autograd.set_detect_anomaly(True)放在训练脚本最开头(比如train.py的main函数第一行),这样能追踪前向传播中所有修改梯度依赖变量的原地操作,输出更详细的异常路径,准确定位触发问题的代码行。 - 排查原地操作场景:
- 检查模型前向传播逻辑,是否存在
+=、*=这类原地赋值,或是带_后缀的PyTorch函数(如relu_),这类操作会直接修改张量本身,破坏梯度计算链。 - 查看数据预处理或batch处理阶段,是否对输入张量做了原地修改,比如直接修改
batch内元素而未创建副本。
- 检查模型前向传播逻辑,是否存在
- 调整混合精度训练写法:若使用
torch.cuda.amp,确保scaler.scale(loss).backward()执行前,loss未被原地修改。可临时验证:先创建loss副本loss_copy = loss.detach().clone(),再用scaler.scale(loss_copy).backward(),但建议优先找到问题根源而非仅用此 workaround。 - 调整PyTorch版本:部分旧版PyTorch在混合精度与原地操作的兼容性上存在bug,比如当前使用1.x版本,可尝试降级到1.12.x或升级到2.x稳定版,验证问题是否消失。
- 检查自定义模块:若模型包含自定义层或函数,重点排查这些部分的张量操作,确保所有需保留梯度的张量未被原地修改,必要时用
.clone()创建副本后再操作。
内容的提问来源于stack exchange,提问作者iniKan
相关产品推荐
相关产品推荐

