You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何transformers.pipeline设device_map='auto'仍仅用CPU?多GPU部署LLaMA-2-13B遇阻

问题:LLaMA-2-13B部署时device_map="auto"未启用多GPU,始终用CPU

我在8×32GB Tesla V100 GPU服务器上部署LLaMA-2-13B,使用transformers.pipeline并设置device_map="auto",但运行脚本时nvidia-smi显示GPU未被占用,始终仅使用CPU。

我的pipeline配置代码:

pipeline = transformers.pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    torch_dtype=torch.float16,
    device_map="auto",
)

sequences = pipeline(
    prompt,
    do_sample=True,
    top_k=10,
    num_return_sequences=1,
    eos_token_id=tokenizer.eos_token_id,
    max_length=512,
)

运行时nvidia-smi输出(GPU显存占用始终为0):

+---------------------------------------------------------------------------------------+
| NVIDIA-SMI 535.129.03             Driver Version: 535.129.03   CUDA Version: 12.2     |
|-----------------------------------------+----------------------+----------------------+| GPU  Name                 Persistence-M | Bus-Id        Disp.A | Volatile Uncorr. ECC || Fan  Temp   Perf          Pwr:Usage/Cap |         Memory-Usage | GPU-Util  Compute M. ||                                         |                      |               MIG M. ||=========================================+======================+======================||   0  Tesla V100-SXM2-32GB           Off | 00000000:06:00.0 Off |                    0 || N/A   35C    P0              45W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   1  Tesla V100-SXM2-32GB           Off | 00000000:07:00.0 Off |                    0 || N/A   35C    P0              43W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   2  Tesla V100-SXM2-32GB           Off | 00000000:0A:00.0 Off |                    0 || N/A   36C    P0              45W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   3  Tesla V100-SXM2-32GB           Off | 00000000:0B:00.0 Off |                    0 || N/A   34C    P0              41W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   4  Tesla V100-SXM2-32GB           Off | 00000000:85:00.0 Off |                    0 || N/A   34C    P0              46W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   5  Tesla V100-SXM2-32GB           Off | 00000000:86:00.0 Off |                    0 || N/A   35C    P0              44W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   6  Tesla V100-SXM2-32GB           Off | 00000000:89:00.0 Off |                    0 || N/A   37C    P0              43W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+|   7  Tesla V100-SXM2-32GB           Off | 00000000:8A:00.0 Off |                    0 || N/A   34C    P0              44W / 300W |      0MiB / 32768MiB |      0%      Default ||                                         |                      |                  N/A |+-----------------------------------------+----------------------+----------------------+

+---------------------------------------------------------------------------------------+
| Processes:                                                                            ||  GPU   GI   CI        PID   Type   Process name                            GPU Memory ||        ID   ID                                                             Usage      ||=======================================================================================||  No running processes found                                                           |+---------------------------------------------------------------------------------------+

尝试指定device="cuda:0"时会使用该GPU,但LLaMA-2-13B所需内存超过单卡32GB,无法运行。希望知道如何设置让pipeline使用多GPU。


解决方案

以下是几个可行的调整步骤:

  • 确保加载模型时就启用device_map
    问题可能出在你已经提前加载了模型到CPU,之后再传入pipeline时,device_map="auto"不会自动迁移模型到GPU。正确的做法是在加载模型时就指定device_map="auto"和torch_dtype=torch.float16:

    from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
    
    tokenizer = AutoTokenizer.from_pretrained("meta-llama/Llama-2-13b-chat-hf")
    model = AutoModelForCausalLM.from_pretrained(
        "meta-llama/Llama-2-13b-chat-hf",
        torch_dtype=torch.float16,
        device_map="auto"
    )
    
    pipeline = transformers.pipeline(
        "text-generation",
        model=model,
        tokenizer=tokenizer,
        torch_dtype=torch.float16
    )
    
  • 安装并启用accelerate库
    device_map功能依赖于accelerate库,确保你已经安装了最新版本:

    pip install --upgrade accelerate
    

    加载模型前可以先初始化accelerate环境,确保多GPU被正确识别:

    from accelerate import Accelerator
    accelerator = Accelerator()
    model = accelerator.prepare(model)
    
  • 检查CUDA可用性与环境变量

    • 运行以下代码确认PyTorch能识别到所有GPU:
      import torch
      print(torch.cuda.device_count())
      print(torch.cuda.is_available())
      
    • 如果输出的设备数为0,说明PyTorch未正确关联CUDA,需要重新安装对应CUDA版本的PyTorch。
    • 确保没有设置CUDA_VISIBLE_DEVICES环境变量限制GPU使用,可通过echo $CUDA_VISIBLE_DEVICES检查。
  • 使用pipeline的model_kwargs参数传递device_map
    另一种方式是在pipeline初始化时,通过model_kwargs传入设备映射配置,确保模型加载阶段就应用多GPU分配:

    pipeline = transformers.pipeline(
        "text-generation",
        model="meta-llama/Llama-2-13b-chat-hf",
        tokenizer="meta-llama/Llama-2-13b-chat-hf",
        model_kwargs={"torch_dtype": torch.float16, "device_map": "auto"}
    )
    

内容的提问来源于stack exchange,提问作者Phil-Antony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 15:45:21