在Paperspace运行Transformers Pipeline调用GPU时遇CUDA内存不足的解决方法
Paperspace中HuggingFace Pipeline启用GPU时CUDA内存不足的解决办法
在Paperspace平台使用GPU运行HuggingFace Transformers Pipeline模型时,设置
device=0出现CUDA内存不足错误:RuntimeError: CUDA out of memory. Tried to allocate 16.00 MiB (GPU 0; 15.90 GiB total capacity; 476.40 MiB already allocated; 7.44 MiB free; 492.00 MiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF不设置
device=0则无法启用GPU,相同模型和代码在Google Colab仅设置device=0即可正常运行,当前代码如下:from transformers import pipeline classifier = pipeline("zero-shot-classification", model="Sahajtomar/German_Zeroshot" #,device = 0 )
以下是几种可行的解决方法:
优化PyTorch内存分配:设置
PYTORCH_CUDA_ALLOC_CONF环境变量减少显存碎片,在代码开头添加:import os os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'max_split_size_mb:128'可根据实际情况调整
max_split_size_mb的值(如64、256),让PyTorch更高效分配显存。启用半精度推理:通过
model_kwargs让模型自动使用半精度加载,大幅降低显存占用,修改后的代码:from transformers import pipeline classifier = pipeline("zero-shot-classification", model="Sahajtomar/German_Zeroshot", device=0, model_kwargs={"torch_dtype": "auto"} )半精度推理对模型效果影响极小,但能节省约50%的显存。
清理残留显存:运行模型前手动清理未释放的显存,避免之前进程占用资源:
import torch torch.cuda.empty_cache()检查显存占用情况:使用
nvidia-smi命令查看GPU显存使用状态,确认是否有其他后台进程占用显存,结束不必要的进程释放资源。
内容的提问来源于stack exchange,提问作者Alexey Bogdanov
相关产品推荐
相关产品推荐

