You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

运行MolScribe训练脚本时遇反向传播原地操作RuntimeError

问题

运行MolScribe仓库中的bash scripts/train_uspto_joint_chartok_1m680k.sh脚本时触发RuntimeError,错误信息如下:

RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.cuda.LongTensor [1, 123]] is at version 1; expected version 0 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).

错误发生在train.py的train_fn函数中scaler.scale(loss).backward()行,完整回溯信息:

Traceback (most recent call last):
  File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 612, in <module>
    main()
  File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 598, in main
    train_loop(args, train_df, valid_df, aux_df, tokenizer, args.save_path)
  File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 380, in train_loop
    avg_loss, global_step = train_fn(
  File "/home/qk/Documents/FTECH/OCR/MolScribe/train.py", line 221, in train_fn
    scaler.scale(loss).backward()
  File "/home/qk/anaconda3/envs/MolScribe/lib/python3.9/site-packages/torch/_tensor.py", line 488, in backward
    torch.autograd.backward(
  File "/home/qk/anaconda3/envs/MolScribe/lib/python3.9/site-packages/torch/autograd/__init__.py", line 197, in backward
    Variable._execution_engine.run_backward(  # Calls into the C++ engine to run the backward pass
RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation: [torch.cuda.LongTensor [1, 123]] is at version 1; expected version 0 instead. Hint: enable anomaly detection to find the operation that failed to compute its gradient, with torch.autograd.set_detect_anomaly(True).

已尝试添加with torch.autograd.set_detect_anomaly(True)调试,但回溯信息无变化,无法定位具体操作。

解决建议

  • 正确启用异常检测:不要仅在backward()附近添加检测,把torch.autograd.set_detect_anomaly(True)放在训练脚本最开头(比如train.py的main函数第一行),这样能追踪前向传播中所有修改梯度依赖变量的原地操作,输出更详细的异常路径,准确定位触发问题的代码行。
  • 排查原地操作场景:
    • 检查模型前向传播逻辑,是否存在+=、*=这类原地赋值,或是带_后缀的PyTorch函数(如relu_),这类操作会直接修改张量本身,破坏梯度计算链。
    • 查看数据预处理或batch处理阶段,是否对输入张量做了原地修改,比如直接修改batch内元素而未创建副本。
  • 调整混合精度训练写法:若使用torch.cuda.amp,确保scaler.scale(loss).backward()执行前,loss未被原地修改。可临时验证:先创建loss副本loss_copy = loss.detach().clone(),再用scaler.scale(loss_copy).backward(),但建议优先找到问题根源而非仅用此 workaround。
  • 调整PyTorch版本:部分旧版PyTorch在混合精度与原地操作的兼容性上存在bug,比如当前使用1.x版本,可尝试降级到1.12.x或升级到2.x稳定版,验证问题是否消失。
  • 检查自定义模块:若模型包含自定义层或函数,重点排查这些部分的张量操作,确保所有需保留梯度的张量未被原地修改,必要时用.clone()创建副本后再操作。

内容的提问来源于stack exchange,提问作者iniKan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 00:37:32