DeepSpeed未按预期向CPU卸载运算,显存不足报错求助
调大batch size训练CLIP模型时触发CUDA OOM错误,DeepSpeed ZeRO Stage2配置的CPU卸载未按预期缓解显存压力,调小batch size则无报错。
报错信息
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 7.04 GiB (GPU 1; 79.15 GiB total capacity; 68.07 GiB already allocated; 5.90 GiB free; 72.14 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF
环境与配置
- 模型:HuggingFace CLIP
- 框架:PyTorch 1.13
- 训练方式:HuggingFace Trainer集成DeepSpeed
- 服务器:Azure虚拟机(AMD EPYC 7V13 64核)
- 优化器:AdamW
- DeepSpeed配置(Stage2):
{ "fp16": { "enabled": "auto", "loss_scale": 0, "loss_scale_window": 1000, "initial_scale_power": 16, "hysteresis": 2, "min_loss_scale": 1 }, "optimizer": { "type": "AdamW", "params": { "lr": "auto", "betas": "auto", "eps": "auto", "weight_decay": "auto" } }, "scheduler": { "type": "WarmupLR", "params": { "warmup_min_lr": "auto", "warmup_max_lr": "auto", "warmup_num_steps": "auto" } }, "zero_optimization": { "stage": 2, "offload_optimizer": { "device": "cpu", "pin_memory": true }, "allgather_partitions": true, "allgather_bucket_size": 2e8, "overlap_comm": true, "reduce_scatter": true, "reduce_bucket_size": 2e8, "contiguous_gradients": true }, "gradient_accumulation_steps": "auto", "gradient_clipping": "auto", "train_batch_size": "auto", "train_micro_batch_size_per_gpu": "auto" }
- Trainer初始化代码:
with open("./Multi_Modal_Model/zero_config/stage_2_config.json") as f: z_optimiser = json.load(f) training_args = TrainingArguments( ... deepspeed=z_optimiser, ... )
1. ZeRO Stage2的卸载范围限制
ZeRO Stage2仅会将优化器状态(如AdamW的动量、方差)卸载到CPU,模型的参数、激活值、梯度仍然保留在GPU上。调大batch size时,模型前向/反向传播产生的激活值显存占用会大幅增加,超过GPU剩余空间就会触发OOM——这是Stage2的设计特性,并非配置错误。
如果需要进一步卸载模型参数/梯度以释放GPU显存,需升级到ZeRO Stage3,修改配置如下:
"zero_optimization": { "stage": 3, "offload_optimizer": { "device": "cpu", "pin_memory": true }, "offload_param": { "device": "cpu", "pin_memory": true }, "allgather_partitions": true, "allgather_bucket_size": 2e8, "overlap_comm": true, "reduce_scatter": true, "reduce_bucket_size": 2e8, "contiguous_gradients": true }
2. 修正"auto"参数的模糊配置
DeepSpeed配置中多个"auto"设置可能导致行为不可控:
- 明确设置
train_micro_batch_size_per_gpu为固定值(与TrainingArguments中的per_device_train_batch_size保持一致),避免自动调整的micro batch导致显存峰值过高。 - 将
fp16.enabled手动设为true,确保混合精度生效(混合精度可将显存占用降低约50%),替代依赖"auto"的自动判断。
3. 缓解显存碎片问题
报错提示reserved内存远大于allocated内存,说明存在显存碎片。设置环境变量缓解:
export PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:512
该设置会限制PyTorch分配的显存块大小,减少碎片产生。
4. 检查HuggingFace Trainer的DeepSpeed集成一致性
- 确保TrainingArguments中的
per_device_train_batch_size与DeepSpeed配置的train_micro_batch_size_per_gpu数值一致,避免参数冲突。 - 查看训练日志,确认DeepSpeed已正确初始化(日志中会出现"DeepSpeed config info"等相关输出),若未初始化需检查Trainer的启动方式(需用
deepspeed命令启动训练脚本,而非直接python)。
内容的提问来源于stack exchange,提问作者paragon00

