You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Mixtral 8x7B双GPU使用pipeline推理异常缓慢问题求助

Mixtral-8x7B双GPU推理慢问题排查与解决

你的问题核心是模型虽加载到GPU,但实际推理时部分/全部计算回退到CPU,导致CPU满载、内存耗尽、GPU闲置。以下是具体原因和解决办法:

问题根源

  1. AutoModelForCausalLM.from_pretrained的device_map="auto"在双GPU场景下可能分配不彻底,部分模型层仍留在CPU
  2. 未启用推理模式,多余的梯度计算占用大量CPU和内存资源
  3. pipeline未明确指定GPU设备,默认逻辑可能回退到CPU执行部分计算

解决步骤与修改后代码

直接替换原代码的模型加载和推理部分,修改后代码如下:

from transformers import pipeline
from transformers import AutoModelForCausalLM, AutoTokenizer, AutoConfig
import time

import torch
from accelerate import init_empty_weights, load_checkpoint_and_dispatch

t1 = time.perf_counter()

model_id = "mistralai/Mixtral-8x7B-Instruct-v0.1"
tokenizer = AutoTokenizer.from_pretrained(model_id)

# 用accelerate精准分配双GPU资源
with init_empty_weights():
    config = AutoConfig.from_pretrained(model_id)
    model = AutoModelForCausalLM.from_config(config)

# 加载模型并分配到双GPU,使用bf16精度减少内存占用
model = load_checkpoint_and_dispatch(
    model,
    model_id,
    device_map="auto",
    dtype=torch.bfloat16,
    no_split_module_classes=["MixtralDecoderLayer"]
)

t2 = time.perf_counter()
print(f"Loading tokenizer and model: took {t2-t1} seconds to execute.")

# 创建pipeline时明确指定GPU设备
code_generator = pipeline(
    'text-generation', 
    model=model, 
    tokenizer=tokenizer,
    device=0  # 指定主GPU,accelerate会自动调度双GPU协同计算
)

t3 = time.perf_counter()
print(f"Creating pipeline: took {t3-t2} seconds to execute.")

# 推理时启用推理模式,关闭梯度计算以节省资源
while True:
    print("\n=========Please type in your question=========================\n")
    user_content = input("\nQuestion: ").strip()
    t1 = time.perf_counter()
    with torch.inference_mode():
        generated_code = code_generator(
            user_content, 
            pad_token_id=tokenizer.eos_token_id, 
            max_new_tokens=20
        )[0]['generated_text']
    t2 = time.perf_counter()
    print(f"Inferencing using the model: took {t2-t1} seconds to execute.")
    print(generated_code)

关键修改说明

  • 使用load_checkpoint_and_dispatch:相比from_pretrained的自动分配,该API更适配多GPU场景,能确保所有模型层都分配到GPU,避免CPU残留计算
  • 启用bf16精度:Mixtral原生支持bf16,可大幅降低GPU内存占用,同时不损失推理精度,提升计算速度
  • 添加torch.inference_mode():彻底关闭梯度计算,消除不必要的CPU和内存开销
  • pipeline指定device:强制pipeline绑定GPU,避免默认逻辑回退到CPU

执行修改后的代码后,再用nvidia-smi查看GPU负载,应该能看到GPU利用率提升,CPU和内存占用下降,推理速度显著加快。

内容的提问来源于stack exchange,提问作者Fifi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 20:21:01