You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在LangChain+VLLM中调用本地LLM实现多GPU推理?

问题描述

尝试使用本地LLM模型进行推理,因需用到8块Quadro RTX 8000多GPU,选择LangChain搭配VLLM(此前用LangChain+Hugging Face Pipeline多GPU时出错且无时间修复)。使用Hugging Face仓库模型时正常,但切换为本地模型路径后,VLLM报错:

does not appear to have a file named config.json. Checkout huggingface repo/None for available files

推测VLLM误将本地路径当作Hugging Face仓库地址查找,部分源码如下:

from fastapi import FastAPI, Request, Form
from fastapi.templating import Jinja2Templates
from fastapi.staticfiles import StaticFiles
import os
from time import time
from langchain.document_loaders import DirectoryLoader, TextLoader
from langchain.text_splitter import CharacterTextSplitter
from langchain.vectorstores import FAISS
from langchain.embeddings import HuggingFaceEmbeddings
from langchain.retrievers.document_compressors import EmbeddingsFilter
from langchain.retrievers import ContextualCompressionRetriever
from langchain.chains import RetrievalQA
import torch
from langchain.llms import VLLM

# load local vector storage
embedding_id = "intfloat/multilingual-e5-large"
docsearch = FAISS.load_local("./faiss_db_{}".format(embedding_id), embeddings)
embeddings_filter = EmbeddingsFilter(embeddings=embeddings, similarity_threshold=0.80)
compression_retriever = ContextualCompressionRetriever(base_compressor=embeddings_filter,
                                                       base_retriever=docsearch.as_retriever())

llm = VLLM(model="/home/account/somewhere/models/model",
           tensor_parallel_size=2,
           trust_remote_code=True,
           max_new_tokens=2048,
           top_k=50,
           top_p=0.01,
           temperature=0.01,
           repetition_penalty=1.5,
           stop=stop_word
)

qa = RetrievalQA.from_chain_type(llm=llm, chain_type="stuff", retriever=compression_retriever)

st = time()
prompt = "questions"
response = qa.run(query=prompt)
et = time()
print(prompt)
print('>', response)
print('>', et-st, 'sec consumed. ')

请问如何在LangChain+VLLM中使用本地模型?或LangChain实现多GPU推理的可行方法?

解决方案

一、修复LangChain+VLLM加载本地模型的问题

1. 确保本地模型文件完整

VLLM要求本地模型目录必须包含config.json、模型权重文件(如pytorch_model.bin或分块权重文件)、tokenizer.json、tokenizer_config.json等核心文件。如果是从Hugging Face下载的模型,需确认所有必要文件已下载完整,避免遗漏。

2. 显式指定本地路径加载参数

初始化VLLM时,添加download_dir参数并设置为本地模型目录,同时确保model参数直接指向本地路径:

llm = VLLM(
    model="/home/account/somewhere/models/model",
    tensor_parallel_size=8,  # 匹配8块GPU的配置
    trust_remote_code=True,
    max_new_tokens=2048,
    top_k=50,
    top_p=0.01,
    temperature=0.01,
    repetition_penalty=1.5,
    stop=stop_word,
    download_dir="/home/account/somewhere/models/model"  # 显式指定本地目录
)

3. 升级VLLM与LangChain版本

旧版本VLLM处理本地路径可能存在逻辑问题,建议升级到最新稳定版:

pip install --upgrade vllm langchain

二、LangChain多GPU推理替代方案

如果VLLM本地加载问题仍未解决,可尝试以下两种多GPU推理方案:

1. Hugging Face Pipeline + Accelerate

通过accelerate库实现多GPU分布式推理,先配置accelerate环境:

accelerate config

再修改LangChain的LLM初始化代码:

from langchain.llms import HuggingFacePipeline
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch

tokenizer = AutoTokenizer.from_pretrained("/home/account/somewhere/models/model")
model = AutoModelForCausalLM.from_pretrained(
    "/home/account/somewhere/models/model",
    device_map="auto",  # 自动分配模型到多GPU
    load_in_8bit=True,  # 可选,降低显存占用
    trust_remote_code=True
)

pipe = pipeline(
    "text-generation",
    model=model,
    tokenizer=tokenizer,
    max_new_tokens=2048,
    temperature=0.01,
    repetition_penalty=1.5,
    device_map="auto"
)

llm = HuggingFacePipeline(pipeline=pipe)

2. LLaMA.cpp(支持GGUF格式模型)

对于支持GGUF格式的模型,可使用LLaMA.cpp的多GPU张量并行模式,LangChain中通过LlamaCpp类调用:

from langchain.llms import LlamaCpp

llm = LlamaCpp(
    model_path="/home/account/somewhere/models/model.gguf",
    n_gpu_layers=-1,  # 将所有模型层分配到GPU
    n_ctx=2048,
    temperature=0.01,
    repetition_penalty=1.5,
    verbose=True
)

内容的提问来源于stack exchange,提问作者배준호

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 00:17:47