You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多GPU运行Mistral-7B-Instruct-v0.2遇显存不足及单GPU占用问题求助

问题描述

我想用本地4张8GB显存的Geforce GTX 1080离线运行Mistral-7B-Instruct-v0.2模型实现对话功能,修改后的脚本如下:

import torch
import json
from transformers import AutoTokenizer, AutoModelForCausalLM

def generate_text(input_text, num_texts=2, max_length=100, num_beams=5, early_stopping=True):
    # Set the GPUs to use
    device_ids = [0, 1, 2, 3]  # Modify this list according to your GPU configuration
    primary_device = f'cuda:{device_ids[0]}'  # Primary device
    torch.cuda.set_device(primary_device)

    # Load the tokenizer and model
    tokenizer = AutoTokenizer.from_pretrained("MistralAI/Mistral-7B-Instruct-v0.2")
    tokenizer.add_special_tokens({'pad_token': '[PAD]'})
    model = AutoModelForCausalLM.from_pretrained("MistralAI/Mistral-7B-Instruct-v0.2").to(primary_device)

    # Move model to GPUs
    model = torch.nn.DataParallel(model, device_ids=device_ids)

    # Tokenize the input text and move to the primary device
    inputs = tokenizer(input_text, return_tensors="pt", max_length=512, padding=True, truncation=True)
    inputs = {k: v.to(primary_device) for k, v in inputs.items()}

    # Generate multiple texts using different random seeds
    generated_texts = []
    for i in range(num_texts):
        # Set the random seed for reproducibility
        torch.manual_seed(i)

        # Generate the text using the model
        with torch.no_grad():
            outputs = model.module.generate(**inputs, max_length=max_length, num_beams=num_beams, early_stopping=early_stopping)

        # Decode and add the generated text to the list
        generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
        generated_texts.append(generated_text)

    return generated_texts

if __name__ == "__main__":
    # Set the input text and style
    input_text = "Tell me a story about a dragon and a princess."

    # Generate texts
    generated_texts = generate_text(input_text)

    # Write the generated texts to a JSON file
    with open("generated_texts.json", "w") as f:
        json.dump(generated_texts, f)

报错信息

运行脚本后出现显存不足错误:

Loading checkpoint shards: 100%|████████████████████████████████████████████████████| 3/3 [00:02<00:00,  1.16it/s]
Traceback (most recent call last):
  File "myscript.py", line 44, in <module>
    generated_texts = generate_text(input_text)
  File "myscript.py", line 14, in generate_text
    model = AutoModelForCausalLM.from_pretrained("MistralAI/Mistral-7B-Instruct-v0.2").to(primary_device)
  File "/home/user/Transformers/lib/python3.8/site-packages/transformers/modeling_utils.py", line 2556, in to
    return super().to(*args, **kwargs)
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1152, in to
    return self._apply(convert)
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 802, in _apply
    module._apply(fn)
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 802, in _apply
    module._apply(fn)
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 802, in _apply
    module._apply(fn)
  [Previous line repeated 2 more times]
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 825, in _apply
    param_applied = fn(param)
  File "/home/user/Transformers/lib/python3.8/site-packages/torch/nn/modules/module.py", line 1150, in convert
    return t.to(device, dtype if t.is_floating_point() or t.is_complex() else None, non_blocking)
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 224.00 MiB. GPU 0 has a total capacity of 7.92 GiB of which 86.81 MiB is free. Including non-PyTorch memory, this process has 7.12 GiB memory in use. Of the allocated memory 7.02 GiB is allocated by PyTorch, and 1.78 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation.  See documentation for Memory Management  (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)

疑问点

  • 当前脚本用DataParallel但还是单GPU运行,是不是必须让模型完整放入单GPU显存才能用多GPU?
  • Mistral-7B官方要求16GB显存,单卡只有8GB,怎么用4张卡拆分运行?
  • 推理场景需要开SLI吗?训练时不需要,推理会不会不一样?
  • 更大的模型比如Mistral-8X7B-v0.1需要100GB显存,单A100 80GB装不下,该怎么处理?
  • 已经设置export CUDA_VISIBLE_DEVICES=0,1,2,3,但问题没解决。

解决方案

1. 核心问题:DataParallel不适合模型分片

你用的DataParallel是数据并行,它要求模型完整加载到主GPU,再复制到其他GPU,所以主GPU必须能装下整个模型,这就是你显存不足的原因。要解决单卡装不下的问题,得用模型并行,把模型的不同层拆分到不同GPU上。

2. 针对Mistral-7B的实现方案

用Hugging Face transformers内置的模型并行功能,直接加载时就分片到多卡:

修改后的脚本

import torch
import json
from transformers import AutoTokenizer, AutoModelForCausalLM

def generate_text(input_text, num_texts=2, max_length=100, num_beams=5, early_stopping=True):
    # 加载tokenizer,用eos_token替代默认缺失的pad_token
    tokenizer = AutoTokenizer.from_pretrained("MistralAI/Mistral-7B-Instruct-v0.2")
    tokenizer.pad_token = tokenizer.eos_token

    # 启用自动模型并行,拆分模型到多卡
    model = AutoModelForCausalLM.from_pretrained(
        "MistralAI/Mistral-7B-Instruct-v0.2",
        device_map="auto",  # 自动分配模型层到可用GPU
        torch_dtype=torch.float16,  # 半精度减少显存占用
        low_cpu_mem_usage=True
    )

    # 处理输入并自动分配到模型所在设备
    inputs = tokenizer(input_text, return_tensors="pt", padding=True, truncation=True, max_length=512)
    inputs = {k: v.to(model.device) for k, v in inputs.items()}

    generated_texts = []
    for i in range(num_texts):
        torch.manual_seed(i)
        with torch.no_grad():
            outputs = model.generate(**inputs, max_length=max_length, num_beams=num_beams, early_stopping=early_stopping)
        generated_text = tokenizer.decode(outputs[0], skip_special_tokens=True)
        generated_texts.append(generated_text)

    return generated_texts

if __name__ == "__main__":
    input_text = "Tell me a story about a dragon and a princess."
    generated_texts = generate_text(input_text)
    with open("generated_texts.json", "w") as f:
        json.dump(generated_texts, f)

关键改动说明

  • device_map="auto":让框架自动把模型的不同层分配到各个GPU,无需手动指定设备,直接解决单卡装不下的问题。
  • torch_dtype=torch.float16:半精度浮点将显存占用减半,Mistral-7B半精度约占13-14GB,4张8GB卡可轻松分摊。
  • 移除了DataParallel和手动设备设置,device_map已自动处理多卡分配逻辑。

3. 关于SLI的问题

不需要开SLI。SLI是NVIDIA针对游戏的多卡协同技术,和AI模型的并行计算逻辑完全无关,AI模型并行靠PyTorch/Hugging Face的框架实现,与SLI无关联。

4. 超大模型(如Mistral-8X7B)的处理方案

如果单A100 80GB装不下,可采用以下两种方案:

  • 模型并行+流水线并行:用accelerate库配置流水线并行,既按层拆分模型到多卡,又拆分输入序列到不同卡处理,进一步降低单卡显存压力。
  • 量化:用BitsAndBytesConfig做4位或8位量化,将模型显存占用再减少75%或50%。比如Mistral-8X7B用4位量化后约占20-25GB,单A100 80GB即可装下,或用两张16GB卡分摊。

量化示例代码

加载模型时添加量化配置:

from transformers import BitsAndBytesConfig

bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_use_double_quant=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16
)

model = AutoModelForCausalLM.from_pretrained(
    "mistralai/Mixtral-8x7B-Instruct-v0.1",
    device_map="auto",
    quantization_config=bnb_config,
    low_cpu_mem_usage=True
)

5. 额外优化建议

  • 关闭其他占用GPU的程序,避免显存被占用。
  • 设置环境变量PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True,减少显存碎片,提升显存利用率。
  • 若显存仍紧张,可调小max_length,或用do_sample=True替代num_beams(beam search显存占用更高)。

内容的提问来源于stack exchange,提问作者E C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 00:43:12