You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

训练PIDNet语义分割模型时遭遇CUDA非法内存访问错误求助

自定义数据集训练PIDNet时遭遇CUDA非法内存访问错误

我在自定义数据集上训练语义分割模型PIDNet,本地PC训练时出现如下错误:

File "tools/train.py", line 223, in <module>
    main()
  File "tools/train.py", line 185, in main
    train(config, epoch, config.TRAIN.END_EPOCH, 
  File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/function.py", line 43, in train
    losses, _, acc, loss_list = model(images, labels, bd_gts)
  File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl
    return forward_call(*input, **kwargs)
  File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/parallel/data_parallel.py", line 169, in forward
    return self.module(*inputs[0], **kwargs[0])
  File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl
    return forward_call(*input, **kwargs)
  File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/utils.py", line 48, in forward
    loss_s = self.sem_loss(outputs[:-1], labels)
  File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl
    return forward_call(*input, **kwargs)
  File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 90, in forward
    return sum([
  File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 91, in <listcomp>
    w * func(x, target)
  File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 72, in _ohem_forward
    pred, ind = pred.contiguous().view(-1,)[mask].contiguous().sort()
RuntimeError: CUDA error: an illegal memory access was encountered
CUDA kernel errors might be asynchronously reported at some other API call,so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1.

我的GPU参数如下:

+-----------------------------------------------------------------------------+
| NVIDIA-SMI 515.86.01    Driver Version: 515.86.01    CUDA Version: 11.7     |
|-------------------------------+----------------------+----------------------+| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC || Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. ||                               |                      |               MIG M. ||===============================+======================+======================||   0  NVIDIA GeForce ...  On   | 00000000:01:00.0  On |                  N/A || 25%   36C    P0    29W / 120W |    648MiB /  6144MiB |      0%      Default ||                               |                      |                  N/A |+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+| Processes:                                                                  ||  GPU   GI   CI        PID   Type   Process name                  GPU Memory ||        ID   ID                                                   Usage      ||=============================================================================||    0   N/A  N/A       902      G   /usr/lib/xorg/Xorg                245MiB ||    0   N/A  N/A      1234      G   /usr/bin/kwin_x11                 123MiB ||    0   N/A  N/A      1289      G   /usr/bin/plasmashell               48MiB ||    0   N/A  N/A      1481      G   /usr/lib/firefox/firefox          173MiB ||    0   N/A  N/A      5801      G   ...RendererForSitePerProcess       49MiB |+-----------------------------------------------------------------------------+

使用的PyTorch版本:

>>> print(torch.__version__)
1.13.1
>>> 

排查与解决建议

  • 按错误提示设置环境变量CUDA_LAUNCH_BLOCKING=1后重新运行,获取精准的错误栈,定位具体出错步骤
  • 检查自定义数据集标签:确认标签值未超出模型定义的类别数,无NaN或异常值,OHEM损失计算时异常mask会触发内存访问错误
  • 降低batch size:即使显存剩余充足,过大的batch size也可能引发CUDA内存访问问题,尝试减半当前batch size
  • 校验数据维度:确认images、labels、bd_gts的形状、维度完全匹配,比如label尺寸需和模型输出特征图尺寸一致
  • 验证CUDA环境兼容性:PyTorch 1.13.1与CUDA 11.7兼容,但可尝试更新显卡驱动至稳定版本,或重新安装PyTorch确保CUDA toolkit配置正确
  • 关闭DataParallel:单GPU训练时注释掉DataParallel相关代码,改用单卡训练测试是否仍报错

内容的提问来源于stack exchange,提问作者programmer_04_03

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:47:56