You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DeepSpeed未按预期向CPU卸载运算,显存不足报错求助

问题描述

调大batch size训练CLIP模型时触发CUDA OOM错误,DeepSpeed ZeRO Stage2配置的CPU卸载未按预期缓解显存压力,调小batch size则无报错。

报错信息

torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 7.04 GiB (GPU 1; 79.15 GiB total capacity; 68.07 GiB already allocated; 5.90 GiB free; 72.14 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

环境与配置

  • 模型:HuggingFace CLIP
  • 框架:PyTorch 1.13
  • 训练方式:HuggingFace Trainer集成DeepSpeed
  • 服务器:Azure虚拟机(AMD EPYC 7V13 64核)
  • 优化器:AdamW
  • DeepSpeed配置(Stage2):
{
    "fp16": {
        "enabled": "auto",
        "loss_scale": 0,
        "loss_scale_window": 1000,
        "initial_scale_power": 16,
        "hysteresis": 2,
        "min_loss_scale": 1
    },

    "optimizer": {
        "type": "AdamW",
        "params": {
            "lr": "auto",
            "betas": "auto",
            "eps": "auto",
            "weight_decay": "auto"
        }
    },

    "scheduler": {
        "type": "WarmupLR",
        "params": {
            "warmup_min_lr": "auto",
            "warmup_max_lr": "auto",
            "warmup_num_steps": "auto"
        }
    },

    "zero_optimization": {
        "stage": 2,
        "offload_optimizer": {
            "device": "cpu",
            "pin_memory": true
        },
        "allgather_partitions": true,
        "allgather_bucket_size": 2e8,
        "overlap_comm": true,
        "reduce_scatter": true,
        "reduce_bucket_size": 2e8,
        "contiguous_gradients": true
    },

    "gradient_accumulation_steps": "auto",
    "gradient_clipping": "auto",
    "train_batch_size": "auto",
    "train_micro_batch_size_per_gpu": "auto"
}
  • Trainer初始化代码:
with open("./Multi_Modal_Model/zero_config/stage_2_config.json") as f:
    z_optimiser = json.load(f)
        
training_args = TrainingArguments(
    ...
    deepspeed=z_optimiser,
    ...
)

问题分析与解决方案

1. ZeRO Stage2的卸载范围限制

ZeRO Stage2仅会将优化器状态(如AdamW的动量、方差)卸载到CPU,模型的参数、激活值、梯度仍然保留在GPU上。调大batch size时,模型前向/反向传播产生的激活值显存占用会大幅增加,超过GPU剩余空间就会触发OOM——这是Stage2的设计特性,并非配置错误。

如果需要进一步卸载模型参数/梯度以释放GPU显存,需升级到ZeRO Stage3,修改配置如下:

"zero_optimization": {
    "stage": 3,
    "offload_optimizer": {
        "device": "cpu",
        "pin_memory": true
    },
    "offload_param": {
        "device": "cpu",
        "pin_memory": true
    },
    "allgather_partitions": true,
    "allgather_bucket_size": 2e8,
    "overlap_comm": true,
    "reduce_scatter": true,
    "reduce_bucket_size": 2e8,
    "contiguous_gradients": true
}

2. 修正"auto"参数的模糊配置

DeepSpeed配置中多个"auto"设置可能导致行为不可控:

  • 明确设置train_micro_batch_size_per_gpu为固定值(与TrainingArguments中的per_device_train_batch_size保持一致),避免自动调整的micro batch导致显存峰值过高。
  • 将fp16.enabled手动设为true,确保混合精度生效(混合精度可将显存占用降低约50%),替代依赖"auto"的自动判断。

3. 缓解显存碎片问题

报错提示reserved内存远大于allocated内存,说明存在显存碎片。设置环境变量缓解:

export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512

该设置会限制PyTorch分配的显存块大小,减少碎片产生。

4. 检查HuggingFace Trainer的DeepSpeed集成一致性

  • 确保TrainingArguments中的per_device_train_batch_size与DeepSpeed配置的train_micro_batch_size_per_gpu数值一致,避免参数冲突。
  • 查看训练日志,确认DeepSpeed已正确初始化(日志中会出现"DeepSpeed config info"等相关输出),若未初始化需检查Trainer的启动方式(需用deepspeed命令启动训练脚本,而非直接python)。

内容的提问来源于stack exchange,提问作者paragon00

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 05:47:36