You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置PyTorch Llama2-13B模型利用双RTX3080 GPU?

问题

拥有2块NVIDIA GeForce RTX 3080 GPU,运行Llama-2-13B-hf模型时仅使用其中一块,出现CUDA显存不足错误。已使用transformers pipeline并设置device_map="auto",相关代码如下:

import locale
locale.getpreferredencoding = lambda: "UTF-8"

import warnings
warnings.filterwarnings("ignore", category=UserWarning, message="You seem to be using the pipelines sequentially on GPU.")

# %pip install accelerate einops langchain -q
# %pip install --upgrade transformers==4.31.0 -q
# %pip install numpy -q
# %pip install scipy -q
# %pip install pandas -q

import torch
import transformers

from transformers import AutoTokenizer

import pandas as pd
import random
import re

torch.cuda.current_device() # 0
torch.cuda.device_count() # 2
torch.cuda.get_device_name(0) # 'NVIDIA GeForce RTX 3080'

device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
print('Using device:', device) # Using device: cuda

import torch.nn as nn
from transformers import LlamaForCausalLM, LlamaTokenizer, pipeline

model_name = "Llama-2-13b-hf"

tokenizer = LlamaTokenizer.from_pretrained(model_name)

pipeline = pipeline(
    "text-generation",
    model=model_name, 
    torch_dtype=torch.float16,
    device_map="auto",
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.eos_token_id,
)

调用pipeline生成文本的代码:

sequences.append(
    pipeline(
        prompt,
        top_k=10,
        max_new_tokens=150,
        num_return_sequences=1,
    )
)

for sequence in sequences:
    generated_text = sequence[0]['generated_text']
    experiment_result.append(generated_text)

运行时出现错误:

\lib\site-packages\accelerate\utils\modeling.py:317, in set_module_tensor_to_device(module, tensor_name, device, value, dtype, fp16_statistics)
315             module._parameters[tensor_name] = param_cls(new_value, requires_grad=old_value.requires_grad)
316 elif isinstance(value, torch.Tensor):
--> 317     new_value = value.to(device)
318 else:
319     new_value = torch.tensor(value, device=device)

RuntimeError: CUDA out of memory. Tried to allocate 136.00 MiB (GPU 0; 10.00 GiB total capacity; 9.05 GiB already allocated; 0 bytes free; 9.24 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.  See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF

解决方案

  • 更换设备分配策略:device_map="auto"会优先将模型层填满单GPU,改用device_map="balanced"可以更均匀地把模型参数分配到多块GPU上,平衡显存占用。修改pipeline初始化代码:
pipeline = pipeline(
    "text-generation",
    model=model_name, 
    torch_dtype=torch.float16,
    device_map="balanced",  # 替换auto为balanced
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.eos_token_id,
)
  • 启用梯度检查点:通过牺牲少量计算时间来大幅降低显存占用,适合大模型推理。在pipeline中添加gradient_checkpointing=True参数:
pipeline = pipeline(
    "text-generation",
    model=model_name, 
    torch_dtype=torch.float16,
    device_map="balanced",
    gradient_checkpointing=True,  # 新增该参数
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.eos_token_id,
)
  • 手动加载模型并分配设备:避免pipeline自动加载时的潜在问题,显式用accelerate的dispatch_model分配模型到多GPU:
from accelerate import dispatch_model

# 先加载模型
model = LlamaForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16)
# 分配到多GPU
model = dispatch_model(model, device_map="balanced")
# 初始化pipeline
pipeline = pipeline(
    "text-generation",
    model=model, 
    tokenizer=tokenizer,
    torch_dtype=torch.float16,
    eos_token_id=tokenizer.eos_token_id,
    pad_token_id=tokenizer.eos_token_id,
)
  • 优化显存分配策略:设置PyTorch环境变量减少显存碎片,在代码最开头添加:
import os
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128"
  • 验证多GPU使用情况:在模型加载后添加代码,确认两块GPU都被利用:
for i in range(torch.cuda.device_count()):
    print(f"GPU {i} 已分配显存: {torch.cuda.memory_allocated(i)/1024**3:.2f} GB")

内容的提问来源于stack exchange,提问作者Brest

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 16:35:03