如何在LangChain+VLLM中调用本地LLM实现多GPU推理?
尝试使用本地LLM模型进行推理,因需用到8块Quadro RTX 8000多GPU,选择LangChain搭配VLLM(此前用LangChain+Hugging Face Pipeline多GPU时出错且无时间修复)。使用Hugging Face仓库模型时正常,但切换为本地模型路径后,VLLM报错:
does not appear to have a file named config.json. Checkout huggingface repo/None for available files
推测VLLM误将本地路径当作Hugging Face仓库地址查找,部分源码如下:
from fastapi import FastAPI, Request, Form from fastapi.templating import Jinja2Templates from fastapi.staticfiles import StaticFiles import os from time import time from langchain.document_loaders import DirectoryLoader, TextLoader from langchain.text_splitter import CharacterTextSplitter from langchain.vectorstores import FAISS from langchain.embeddings import HuggingFaceEmbeddings from langchain.retrievers.document_compressors import EmbeddingsFilter from langchain.retrievers import ContextualCompressionRetriever from langchain.chains import RetrievalQA import torch from langchain.llms import VLLM # load local vector storage embedding_id = "intfloat/multilingual-e5-large" docsearch = FAISS.load_local("./faiss_db_{}".format(embedding_id), embeddings) embeddings_filter = EmbeddingsFilter(embeddings=embeddings, similarity_threshold=0.80) compression_retriever = ContextualCompressionRetriever(base_compressor=embeddings_filter, base_retriever=docsearch.as_retriever()) llm = VLLM(model="/home/account/somewhere/models/model", tensor_parallel_size=2, trust_remote_code=True, max_new_tokens=2048, top_k=50, top_p=0.01, temperature=0.01, repetition_penalty=1.5, stop=stop_word ) qa = RetrievalQA.from_chain_type(llm=llm, chain_type="stuff", retriever=compression_retriever) st = time() prompt = "questions" response = qa.run(query=prompt) et = time() print(prompt) print('>', response) print('>', et-st, 'sec consumed. ')
请问如何在LangChain+VLLM中使用本地模型?或LangChain实现多GPU推理的可行方法?
一、修复LangChain+VLLM加载本地模型的问题
1. 确保本地模型文件完整
VLLM要求本地模型目录必须包含config.json、模型权重文件(如pytorch_model.bin或分块权重文件)、tokenizer.json、tokenizer_config.json等核心文件。如果是从Hugging Face下载的模型,需确认所有必要文件已下载完整,避免遗漏。
2. 显式指定本地路径加载参数
初始化VLLM时,添加download_dir参数并设置为本地模型目录,同时确保model参数直接指向本地路径:
llm = VLLM( model="/home/account/somewhere/models/model", tensor_parallel_size=8, # 匹配8块GPU的配置 trust_remote_code=True, max_new_tokens=2048, top_k=50, top_p=0.01, temperature=0.01, repetition_penalty=1.5, stop=stop_word, download_dir="/home/account/somewhere/models/model" # 显式指定本地目录 )
3. 升级VLLM与LangChain版本
旧版本VLLM处理本地路径可能存在逻辑问题,建议升级到最新稳定版:
pip install --upgrade vllm langchain
二、LangChain多GPU推理替代方案
如果VLLM本地加载问题仍未解决,可尝试以下两种多GPU推理方案:
1. Hugging Face Pipeline + Accelerate
通过accelerate库实现多GPU分布式推理,先配置accelerate环境:
accelerate config
再修改LangChain的LLM初始化代码:
from langchain.llms import HuggingFacePipeline from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline import torch tokenizer = AutoTokenizer.from_pretrained("/home/account/somewhere/models/model") model = AutoModelForCausalLM.from_pretrained( "/home/account/somewhere/models/model", device_map="auto", # 自动分配模型到多GPU load_in_8bit=True, # 可选,降低显存占用 trust_remote_code=True ) pipe = pipeline( "text-generation", model=model, tokenizer=tokenizer, max_new_tokens=2048, temperature=0.01, repetition_penalty=1.5, device_map="auto" ) llm = HuggingFacePipeline(pipeline=pipe)
2. LLaMA.cpp(支持GGUF格式模型)
对于支持GGUF格式的模型,可使用LLaMA.cpp的多GPU张量并行模式,LangChain中通过LlamaCpp类调用:
from langchain.llms import LlamaCpp llm = LlamaCpp( model_path="/home/account/somewhere/models/model.gguf", n_gpu_layers=-1, # 将所有模型层分配到GPU n_ctx=2048, temperature=0.01, repetition_penalty=1.5, verbose=True )
内容的提问来源于stack exchange,提问作者배준호

