如何配置PyTorch Llama2-13B模型利用双RTX3080 GPU?
问题
拥有2块NVIDIA GeForce RTX 3080 GPU,运行Llama-2-13B-hf模型时仅使用其中一块,出现CUDA显存不足错误。已使用transformers pipeline并设置device_map="auto",相关代码如下:
import locale locale.getpreferredencoding = lambda: "UTF-8" import warnings warnings.filterwarnings("ignore", category=UserWarning, message="You seem to be using the pipelines sequentially on GPU.") # %pip install accelerate einops langchain -q # %pip install --upgrade transformers==4.31.0 -q # %pip install numpy -q # %pip install scipy -q # %pip install pandas -q import torch import transformers from transformers import AutoTokenizer import pandas as pd import random import re torch.cuda.current_device() # 0 torch.cuda.device_count() # 2 torch.cuda.get_device_name(0) # 'NVIDIA GeForce RTX 3080' device = torch.device('cuda' if torch.cuda.is_available() else 'cpu') print('Using device:', device) # Using device: cuda import torch.nn as nn from transformers import LlamaForCausalLM, LlamaTokenizer, pipeline model_name = "Llama-2-13b-hf" tokenizer = LlamaTokenizer.from_pretrained(model_name) pipeline = pipeline( "text-generation", model=model_name, torch_dtype=torch.float16, device_map="auto", eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, )
调用pipeline生成文本的代码:
sequences.append( pipeline( prompt, top_k=10, max_new_tokens=150, num_return_sequences=1, ) ) for sequence in sequences: generated_text = sequence[0]['generated_text'] experiment_result.append(generated_text)
运行时出现错误:
\lib\site-packages\accelerate\utils\modeling.py:317, in set_module_tensor_to_device(module, tensor_name, device, value, dtype, fp16_statistics) 315 module._parameters[tensor_name] = param_cls(new_value, requires_grad=old_value.requires_grad) 316 elif isinstance(value, torch.Tensor): --> 317 new_value = value.to(device) 318 else: 319 new_value = torch.tensor(value, device=device) RuntimeError: CUDA out of memory. Tried to allocate 136.00 MiB (GPU 0; 10.00 GiB total capacity; 9.05 GiB already allocated; 0 bytes free; 9.24 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation. See documentation for Memory Management and PYTORCH_CUDA_ALLOC_CONF
解决方案
- 更换设备分配策略:
device_map="auto"会优先将模型层填满单GPU,改用device_map="balanced"可以更均匀地把模型参数分配到多块GPU上,平衡显存占用。修改pipeline初始化代码:
pipeline = pipeline( "text-generation", model=model_name, torch_dtype=torch.float16, device_map="balanced", # 替换auto为balanced eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, )
- 启用梯度检查点:通过牺牲少量计算时间来大幅降低显存占用,适合大模型推理。在pipeline中添加
gradient_checkpointing=True参数:
pipeline = pipeline( "text-generation", model=model_name, torch_dtype=torch.float16, device_map="balanced", gradient_checkpointing=True, # 新增该参数 eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, )
- 手动加载模型并分配设备:避免pipeline自动加载时的潜在问题,显式用
accelerate的dispatch_model分配模型到多GPU:
from accelerate import dispatch_model # 先加载模型 model = LlamaForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16) # 分配到多GPU model = dispatch_model(model, device_map="balanced") # 初始化pipeline pipeline = pipeline( "text-generation", model=model, tokenizer=tokenizer, torch_dtype=torch.float16, eos_token_id=tokenizer.eos_token_id, pad_token_id=tokenizer.eos_token_id, )
- 优化显存分配策略:设置PyTorch环境变量减少显存碎片,在代码最开头添加:
import os os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "max_split_size_mb:128"
- 验证多GPU使用情况:在模型加载后添加代码,确认两块GPU都被利用:
for i in range(torch.cuda.device_count()): print(f"GPU {i} 已分配显存: {torch.cuda.memory_allocated(i)/1024**3:.2f} GB")
内容的提问来源于stack exchange,提问作者Brest
相关产品推荐
相关产品推荐

