You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch加载HuggingFace模型在GCP无法正常加载到GPU问题

问题现象

编写如下类代码加载HuggingFace的OPT预训练模型与分词器:

class OPT:
    def __init__(self, model_name: str = "facebook/opt-2.7b", use_gpu: bool = False):
        self.model_name = model_name
        self.use_gpu = use_gpu and torch.cuda.is_available()
        print(f"Use gpu:: {self.use_gpu}")

        if self.use_gpu:
            print("Using gpu")
            self.model = AutoModelForCausalLM.from_pretrained(
                self.model_name, torch_dtype=torch.float16
            ).cuda()
        else:
            print("Using cpu")
            self.model = AutoModelForCausalLM.from_pretrained(
                self.model_name, torch_dtype=torch.float32, low_cpu_mem_usage=True
            )

        # the fast tokenizer currently does not work correctly
        self.tokenizer = AutoTokenizer.from_pretrained(model_name, use_fast=False)
  • 代码在Google Colab环境运行完全正常,模型可以顺利加载到GPU显存
  • 迁移到Google Cloud Platform(GCP)环境后,模型始终无法加载到GPU
  • 运行日志显示已经正确识别可用GPU,进入了GPU加载分支:
Use gpu:: True
Using gpu
  • 执行nvidia-smi查看GPU状态时,GPU显存占用为0,无对应运行进程:
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 470.82.01    Driver Version: 470.82.01    CUDA Version: 11.4     |
|-------------------------------+----------------------+----------------------+
| GPU  Name        Persistence-M| Bus-Id        Disp.A | Volatile Uncorr. ECC |
| Fan  Temp  Perf  Pwr:Usage/Cap|         Memory-Usage | GPU-Util  Compute M. |
|                               |                      |               MIG M. |
|===============================+======================+======================|
|   0  Tesla T4            On   | 00000000:00:04.0 Off |                    0 |
| N/A   40C    P8     9W /  70W |      0MiB / 15109MiB |      0%      Default |
|                               |                      |                  N/A |
+-------------------------------+----------------------+----------------------+
                                                                               
+-----------------------------------------------------------------------------+
| Processes:                                                                  |
|  GPU   GI   CI        PID   Type   Process name                  GPU Memory |
|        ID   ID                                                   Usage      |
|=============================================================================|
|  No running processes found                                                 |
+-----------------------------------------------------------------------------+
  • 通过htop观测系统状态,能看到对应Python进程持续占用CPU内存。
排查&解决方法

按以下顺序逐一排查即可解决:

  1. 优先替换模型加载逻辑,不要手动调用.cuda()
    手动在模型加载完成后调用.cuda()存在权重懒加载不触发的兼容问题,尤其是不同环境transformers、PyTorch版本不一致时很容易出现。直接在from_pretrained阶段指定device_map="auto",让库自动完成设备分配,同时给GPU分支也加上low_cpu_mem_usage=True减少CPU内存峰值,修改后的GPU分支代码如下:
    self.model = AutoModelForCausalLM.from_pretrained(
        self.model_name,
        torch_dtype=torch.float16,
        device_map="auto",
        low_cpu_mem_usage=True
    )
    
  2. 验证PyTorch CUDA功能的可用性
    不要只看torch.cuda.is_available()的返回值,在加载模型前先执行一行测试代码验证CUDA能否实际工作:
    print(torch.randn(2, 2).cuda())
    
    如果这行代码报错、卡住超过10秒没有输出,说明你GCP环境安装的PyTorch CUDA版本和宿主机470驱动(对应最高支持CUDA11.4)不兼容,重装匹配CUDA11.4版本的PyTorch即可。
  3. 检查GPU设备访问权限
    部分GCP自定义镜像默认限制非root用户访问NVIDIA设备,先用sudo权限运行你的脚本测试,如果sudo下模型能正常加载到GPU,给当前运行用户添加video、render用户组即可解决权限问题。
  4. 执行一次前向调用再观测显存
    部分版本的PyTorch会延迟CUDA上下文初始化,加载完模型后如果没有执行任何实际的GPU算子调用,不会真正把权重复制到显存,也不会在nvidia-smi里显示进程。可以在模型初始化完成后跑一次简单的生成测试,再看显存占用:
    inputs = tokenizer("test", return_tensors="pt").to("cuda")
    out = model.generate(**inputs, max_new_tokens=10)
    

内容的提问来源于stack exchange,提问作者Nazareno De Francesco

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.30 23:24:14