You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

明明有充足CUDA显存却报内存不足,问题出在哪?

问题:CUDA显存充足但训练时仍报内存不足

硬件配置

!nvidia-smi
Tue Nov 15 08:49:04 2022       
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 510.60.02    Driver Version: 510.60.02    CUDA Version: 11.6     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Quadro RTX 4000     On   | 00000000:81:00.0 Off |                  N/A |
| 44%   32C    P8     9W / 125W |    159MiB /  8192MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|    0   N/A  N/A      2063      G                                      63MiB |
|    0   N/A  N/A   1849271      C                                      91MiB |
+-----------------------------------------------------------------------------+

!free -h
              total        used        free      shared  buff/cache   available
Mem:            64G        677M         31G         10M         32G         63G
Swap:            0B          0B          0B

从nvidia-smi结果看,GPU显存仅占用159MiB,剩余空间充足,但运行PyTorch Lightning训练代码时出现以下内存不足错误:

错误日志

Traceback (most recent call last):
  File "main.py", line 834, in <module>
    raise err
  File "main.py", line 816, in <module>
    trainer.fit(model, data)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py", line 771, in fit
    self._call_and_handle_interrupt(
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py", line 722, in _call_and_handle_interrupt
    return self.strategy.launcher.launch(trainer_fn, *args, trainer=self, **kwargs)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/strategies/launchers/subprocess_script.py", line 93, in launch
    return function(*args, **kwargs)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py", line 812, in _fit_impl
    results = self._run(model, ckpt_path=self.ckpt_path)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/trainer/trainer.py", line 1218, in _run
    self.strategy.setup(self)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/strategies/ddp.py", line 162, in setup
    self.model_to_device()
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/strategies/ddp.py", line 324, in model_to_device
    self.model.to(self.root_device)
  File "/opt/conda/lib/python3.8/site-packages/pytorch_lightning/core/mixins/device_dtype_mixin.py", line 121, in to
    return super().to(*args, **kwargs)
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 927, in to
    return self._apply(convert)
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 579, in _apply
    module._apply(fn)
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 579, in _apply
    module._apply(fn)
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 579, in _apply
    module._apply(fn)
  [Previous line repeated 4 more times]
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 602, in _apply
    param_applied = fn(param)
  File "/opt/conda/lib/python3.8/site-packages/torch/nn/modules/module.py", line 925, in convert
    return t.to(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking)
RuntimeError: CUDA out of memory. Tried to allocate 20.00 MiB (GPU 0; 7.80 GiB total capacity; 6.70 GiB already allocated; 12.44 MiB free; 6.80 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

尝试的解决方法

用以下代码清理GPU缓存后,报错中的"Tried to allocate"数值从146MB降至20MB,但问题依旧:

import torch
from GPUtil import showUtilization as gpu_usage
from numba import cuda

def free_gpu_cache():
    print("Initial GPU Usage")
    gpu_usage()                             

    torch.cuda.empty_cache()

    cuda.select_device(0)
    cuda.close()
    cuda.select_device(0)

    print("GPU Usage after emptying the cache")
    gpu_usage()

free_gpu_cache()

请问哪里操作错了?


问题分析与解决建议

核心问题:显存统计差异的本质

nvidia-smi显示的是系统层面已分配的显存,而PyTorch报错里的是进程内部的显存池状态。你看到的低占用是因为PyTorch默认会预分配大块显存作为缓存池(报错里的6.80 GiB reserved),但实际使用后剩余的缓存空间已经碎片化,连20MiB的连续显存块都无法分配,所以触发OOM。

当前清理方法的无效原因

torch.cuda.empty_cache()只能释放PyTorch中已废弃但未归还系统的显存,不会触动预分配的缓存池,因此无法解决碎片化问题。另外混用numba.cuda的设备操作属于多余步骤,反而可能干扰PyTorch的显存管理逻辑。

实际解决步骤

  1. 关闭预分配,改用按需显存分配
    启动代码前设置环境变量,让PyTorch仅在需要时分配显存,避免预占大块空间:

    export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
    

    或在Python代码开头添加:

    import os
    os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'expandable_segments:True'
    
  2. 缓解显存碎片化
    按照报错提示,设置max_split_size_mb参数,强制PyTorch在分配时避免过度碎片化:

    export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:128
    

    可根据模型调整数值(如64、256),找到适配的参数。

  3. 检查DDP模式的显存占用
    你使用的是DDP策略,注意DDP会为每个进程加载一份模型副本。若脚本误启动多进程(如未正确设置devices参数),会导致显存被多份模型占满。可先切换到单GPU模式验证:

    trainer = Trainer(accelerator='gpu', devices=1, strategy='single_device')
    
  4. 模型与数据加载优化

    • 降低batch_size,或用gradient_accumulation_steps模拟大批次训练
    • 启用半精度训练(precision=16),PyTorch Lightning直接支持该参数
    • 检查数据加载器是否提前将数据放到GPU,造成额外显存占用

内容的提问来源于stack exchange,提问作者mchd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 15:50:21