You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练LayoutLMv3触发CUDA device-side assert错误,求解决(V100 4卡)

LayoutLMv3训练时CUDA设备断言触发错误的解决方法

错误信息

执行trainer.train()启动训练时出现以下错误:

RuntimeError                              Traceback (most recent call last)
/tmp/ipykernel_3844/4032920361.py in <module>
----> 1 trainer.train()

/data/anaconda3/envs/data/lib/python3.7/site-packages/transformers/trainer.py in train(self, resume_from_checkpoint, trial, ignore_keys_for_eval, **kwargs)
   1417             resume_from_checkpoint=resume_from_checkpoint,
   1418             trial=trial,
-> 1419             ignore_keys_for_eval=ignore_keys_for_eval,
   1420         )
   1421 

/data/anaconda3/envs/data/lib/python3.7/site-packages/transformers/trainer.py in _inner_training_loop(self, batch_size, args, resume_from_checkpoint, trial, ignore_keys_for_eval)
   1655                         tr_loss_step = self.training_step(model, inputs)
   1656                 else:
-> 1657                     tr_loss_step = self.training_step(model, inputs)
   1658 
   1659                 if (

/data/anaconda3/envs/data/lib/python3.7/site-packages/transformers/trainer.py in training_step(self, model, inputs)
   2348 
   2349         with self.compute_loss_context_manager():
-> 2350             loss = self.compute_loss(model, inputs)
   2351 
   2352         if self.args.n_gpu > 1:
...
    visual_bbox = visual_bbox.to(device).type(dtype)
RuntimeError: CUDA error: device-side assert triggered
CUDA kernel errors might be asynchronously reported at some other API call,so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.

排查与解决步骤

1. 启用同步CUDA调试定位真实错误

按提示设置环境变量强制CUDA同步执行,让错误直接触发在实际出错位置,获取准确报错栈:

  • 启动训练前执行:
    export CUDA_LAUNCH_BLOCKING=1
    
    或者在Python代码开头添加:
    import os
    os.environ['CUDA_LAUNCH_BLOCKING'] = '1'
    
    重新运行训练,定位具体出错环节。

2. 检查Bounding Box数据合法性

错误出现在bbox转设备的步骤,重点验证数据格式:

  • 确认所有bbox为[x0, y0, x1, y1]格式,且满足0 ≤ x0 ≤ x1 ≤ 图像宽度、0 ≤ y0 ≤ y1 ≤ 图像高度,无负数、超界值或x0>x1、y0>y1的异常。
  • 过滤或修复含NaN、全0等无效值的样本。

3. 多GPU设备一致性检查

确保4块V100上的数据分发正常:

  • 验证Trainer参数中设备分配逻辑,避免数据滞留CPU未同步到GPU。
  • 检查数据加载器的collate_fn,确保多GPU场景下张量拼接无形状不匹配问题。

4. 模型与数据类型匹配

确认数据类型和模型参数一致:

  • 若使用混合精度训练(如float16),确保bbox数据转换为对应类型,避免类型不兼容。
  • 检查visual_bbox原始数据类型,避免数值范围超出目标类型表示范围。

5. 样本序列长度校验

检查输入文本、bbox序列长度是否超过模型max_seq_length上限:

  • 确保过长样本已正确截断,避免张量维度不匹配触发断言错误。

内容的提问来源于stack exchange,提问作者rusubbiz muzkaq

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 09:06:18