You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Paperspace运行Transformers Pipeline调用GPU时遇CUDA内存不足的解决方法

Paperspace中HuggingFace Pipeline启用GPU时CUDA内存不足的解决办法

在Paperspace平台使用GPU运行HuggingFace Transformers Pipeline模型时,设置device=0出现CUDA内存不足错误:

RuntimeError: CUDA out of memory. Tried to allocate 16.00 MiB (GPU 0; 15.90 GiB total capacity; 476.40 MiB already allocated; 7.44 MiB free; 492.00 MiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

不设置device=0则无法启用GPU,相同模型和代码在Google Colab仅设置device=0即可正常运行,当前代码如下:

from transformers import pipeline
classifier = pipeline("zero-shot-classification",
                      model="Sahajtomar/German_Zeroshot"
                     #,device = 0
                     )

以下是几种可行的解决方法:

  • 优化PyTorch内存分配:设置PYTORCH_CUDA_ALLOC_CONF环境变量减少显存碎片,在代码开头添加:

    import os
    os.environ['PYTORCH_CUDA_ALLOC_CONF'] = 'max_split_size_mb:128'
    

    可根据实际情况调整max_split_size_mb的值(如64、256),让PyTorch更高效分配显存。

  • 启用半精度推理:通过model_kwargs让模型自动使用半精度加载,大幅降低显存占用,修改后的代码:

    from transformers import pipeline
    classifier = pipeline("zero-shot-classification",
                          model="Sahajtomar/German_Zeroshot",
                          device=0,
                          model_kwargs={"torch_dtype": "auto"}
                         )
    

    半精度推理对模型效果影响极小,但能节省约50%的显存。

  • 清理残留显存:运行模型前手动清理未释放的显存,避免之前进程占用资源:

    import torch
    torch.cuda.empty_cache()
    
  • 检查显存占用情况:使用nvidia-smi命令查看GPU显存使用状态,确认是否有其他后台进程占用显存,结束不必要的进程释放资源。

内容的提问来源于stack exchange,提问作者Alexey Bogdanov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 19:57:25