You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LLM文本生成训练遇torch.cuda.OutOfMemoryError求助

解决CUDA OutOfMemoryError(剩余显存充足仍报错)的方案

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 16.00 MiB. GPU 0 has a total capacty of 6.00 GiB of which 4.54 GiB is free. Of the allocated memory 480.02 MiB is allocated by PyTorch, and 1.98 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

  • 优化PyTorch内存分配:按错误提示调整max_split_size_mb参数,解决显存碎片化问题。
    启动训练前执行环境变量设置:

    export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128
    

    或者在Python代码开头添加:

    import os
    os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'max_split_size_mb:128'
    

    可以尝试64、256等不同数值,找到适配你场景的参数。

  • 清理显存碎片:在训练循环里定期调用torch.cuda.empty_cache(),比如每个epoch结束后执行一次,不要太频繁以免影响性能。

  • 排查隐性显存占用:

    • 用nvidia-smi命令查看GPU进程,杀掉无关进程释放显存。
    • 检查代码里有没有未释放的大Tensor,比如全局变量里的冗余张量、没删除的中间计算结果。
  • 降低训练负载:

    • 直接减小batch size,这是最有效的显存优化手段。
    • 开启混合精度训练,用torch.cuda.amp.GradScaler()和torch.autocast(device_type='cuda')包裹训练步骤,大幅减少显存占用。
    • 微调LLM的话,用LoRA(低秩适配)方法,只训练部分参数,显存需求会骤降。
  • 升级PyTorch版本:老版本PyTorch可能存在显存分配bug,升级到2.0以上的稳定版大概率能解决碎片化问题。

内容的提问来源于stack exchange,提问作者Dre174

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 12:30:01