训练PIDNet语义分割模型时遭遇CUDA非法内存访问错误求助
自定义数据集训练PIDNet时遭遇CUDA非法内存访问错误
我在自定义数据集上训练语义分割模型PIDNet,本地PC训练时出现如下错误:
File "tools/train.py", line 223, in <module> main() File "tools/train.py", line 185, in main train(config, epoch, config.TRAIN.END_EPOCH, File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/function.py", line 43, in train losses, _, acc, loss_list = model(images, labels, bd_gts) File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl return forward_call(*input, **kwargs) File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/parallel/data_parallel.py", line 169, in forward return self.module(*inputs[0], **kwargs[0]) File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl return forward_call(*input, **kwargs) File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/utils.py", line 48, in forward loss_s = self.sem_loss(outputs[:-1], labels) File "/home/deshpand/anaconda3/envs/torch_env/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1194, in _call_impl return forward_call(*input, **kwargs) File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 90, in forward return sum([ File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 91, in <listcomp> w * func(x, target) File "/home/deshpand/thesis_rr/semantic_segmentation_network/PIDNet/tools/../utils/criterion.py", line 72, in _ohem_forward pred, ind = pred.contiguous().view(-1,)[mask].contiguous().sort() RuntimeError: CUDA error: an illegal memory access was encountered CUDA kernel errors might be asynchronously reported at some other API call,so the stacktrace below might be incorrect. For debugging consider passing CUDA_LAUNCH_BLOCKING=1.
我的GPU参数如下:
+-----------------------------------------------------------------------------+ | NVIDIA-SMI 515.86.01 Driver Version: 515.86.01 CUDA Version: 11.7 | |-------------------------------+----------------------+----------------------+| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC || Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. || | | MIG M. ||===============================+======================+======================|| 0 NVIDIA GeForce ... On | 00000000:01:00.0 On | N/A || 25% 36C P0 29W / 120W | 648MiB / 6144MiB | 0% Default || | | N/A |+-------------------------------+----------------------+----------------------+ +-----------------------------------------------------------------------------+| Processes: || GPU GI CI PID Type Process name GPU Memory || ID ID Usage ||=============================================================================|| 0 N/A N/A 902 G /usr/lib/xorg/Xorg 245MiB || 0 N/A N/A 1234 G /usr/bin/kwin_x11 123MiB || 0 N/A N/A 1289 G /usr/bin/plasmashell 48MiB || 0 N/A N/A 1481 G /usr/lib/firefox/firefox 173MiB || 0 N/A N/A 5801 G ...RendererForSitePerProcess 49MiB |+-----------------------------------------------------------------------------+
使用的PyTorch版本:
>>> print(torch.__version__) 1.13.1 >>>
排查与解决建议
- 按错误提示设置环境变量
CUDA_LAUNCH_BLOCKING=1后重新运行,获取精准的错误栈,定位具体出错步骤 - 检查自定义数据集标签:确认标签值未超出模型定义的类别数,无NaN或异常值,OHEM损失计算时异常mask会触发内存访问错误
- 降低batch size:即使显存剩余充足,过大的batch size也可能引发CUDA内存访问问题,尝试减半当前batch size
- 校验数据维度:确认images、labels、bd_gts的形状、维度完全匹配,比如label尺寸需和模型输出特征图尺寸一致
- 验证CUDA环境兼容性:PyTorch 1.13.1与CUDA 11.7兼容,但可尝试更新显卡驱动至稳定版本,或重新安装PyTorch确保CUDA toolkit配置正确
- 关闭DataParallel:单GPU训练时注释掉DataParallel相关代码,改用单卡训练测试是否仍报错
内容的提问来源于stack exchange,提问作者programmer_04_03
相关产品推荐
相关产品推荐

