24GB显存GPU训练GPT-2时CUDA内存不足问题求助
OutOfMemoryError: CUDA out of memory. Tried to allocate 64.00 MiB (GPU 0; 23.68 GiB total capacity; 18.17 GiB already allocated; 64.62 MiB free; 18.60 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF.
环境配置与训练参数
- GPU:配备24GB VRAM的NVIDIA GPU
- 模型:GPT-2,大小约3GB,含800个32位参数
- 训练数据:36000条训练样本,向量长度600
- 训练配置:5个epoch,batch size为16,启用fp16
内存占用计算
模型大小
GPT-2 model: ~3 GB
梯度
Gradients are typically of the same size as the model’s parameters.
批次大小与训练样本
Batch Size: 16
Training Examples: 36,000
Vector Length: 600
单批次内存分配:
Model: 3 GB (每批次无变化)
Gradients: 3 GB (每批次无变化)
输入数据:16 × 600(向量长度)× 4字节(假设为32位浮点型)= 每批次37.5 KB
输出数据:16 × 600(向量长度)× 4字节(假设为32位浮点型)= 每批次37.5 KB
根据上述计算,该场景中单批次内存分配约为:
- Model: 3 GB
- Gradients: 3 GB
- 输入与输出数据:75 KB
恳请提供问题排查思路或解决建议,感谢您的帮助!
内容的提问来源于stack exchange,提问作者Humza Sami

