Mixtral 8x7B双GPU使用pipeline推理异常缓慢问题求助
Mixtral-8x7B双GPU推理慢问题排查与解决
你的问题核心是模型虽加载到GPU,但实际推理时部分/全部计算回退到CPU,导致CPU满载、内存耗尽、GPU闲置。以下是具体原因和解决办法:
问题根源
AutoModelForCausalLM.from_pretrained的device_map="auto"在双GPU场景下可能分配不彻底,部分模型层仍留在CPU- 未启用推理模式,多余的梯度计算占用大量CPU和内存资源
- pipeline未明确指定GPU设备,默认逻辑可能回退到CPU执行部分计算
解决步骤与修改后代码
直接替换原代码的模型加载和推理部分,修改后代码如下:
from transformers import pipeline from transformers import AutoModelForCausalLM, AutoTokenizer, AutoConfig import time import torch from accelerate import init_empty_weights, load_checkpoint_and_dispatch t1 = time.perf_counter() model_id = "mistralai/Mixtral-8x7B-Instruct-v0.1" tokenizer = AutoTokenizer.from_pretrained(model_id) # 用accelerate精准分配双GPU资源 with init_empty_weights(): config = AutoConfig.from_pretrained(model_id) model = AutoModelForCausalLM.from_config(config) # 加载模型并分配到双GPU,使用bf16精度减少内存占用 model = load_checkpoint_and_dispatch( model, model_id, device_map="auto", dtype=torch.bfloat16, no_split_module_classes=["MixtralDecoderLayer"] ) t2 = time.perf_counter() print(f"Loading tokenizer and model: took {t2-t1} seconds to execute.") # 创建pipeline时明确指定GPU设备 code_generator = pipeline( 'text-generation', model=model, tokenizer=tokenizer, device=0 # 指定主GPU,accelerate会自动调度双GPU协同计算 ) t3 = time.perf_counter() print(f"Creating pipeline: took {t3-t2} seconds to execute.") # 推理时启用推理模式,关闭梯度计算以节省资源 while True: print("\n=========Please type in your question=========================\n") user_content = input("\nQuestion: ").strip() t1 = time.perf_counter() with torch.inference_mode(): generated_code = code_generator( user_content, pad_token_id=tokenizer.eos_token_id, max_new_tokens=20 )[0]['generated_text'] t2 = time.perf_counter() print(f"Inferencing using the model: took {t2-t1} seconds to execute.") print(generated_code)
关键修改说明
- 使用
load_checkpoint_and_dispatch:相比from_pretrained的自动分配,该API更适配多GPU场景,能确保所有模型层都分配到GPU,避免CPU残留计算 - 启用bf16精度:Mixtral原生支持bf16,可大幅降低GPU内存占用,同时不损失推理精度,提升计算速度
- 添加
torch.inference_mode():彻底关闭梯度计算,消除不必要的CPU和内存开销 - pipeline指定
device:强制pipeline绑定GPU,避免默认逻辑回退到CPU
执行修改后的代码后,再用nvidia-smi查看GPU负载,应该能看到GPU利用率提升,CPU和内存占用下降,推理速度显著加快。
内容的提问来源于stack exchange,提问作者Fifi
相关产品推荐
相关产品推荐

