如何基于FastAPI、Llama.cpp与LangChain实现本地LLM流式响应?
FastAPI + Llama.cpp + LangChain 逐Token流式响应实现方案
核心思路
要实现FastAPI的流式响应,需结合LangChain的回调机制捕获逐token输出,再通过FastAPI的StreamingResponse将token实时返回给客户端。以下是完整可运行的修改方案:
完整代码实现
from fastapi import FastAPI, StreamingResponse from langchain.llms import LlamaCpp from langchain.callbacks.base import BaseCallbackHandler from langchain.prompts import PromptTemplate import asyncio from queue import Queue from typing import Any import threading app = FastAPI() # 自定义流式回调处理器:捕获LLM生成的每个token并存入队列 class StreamingCallbackHandler(BaseCallbackHandler): def __init__(self, queue: Queue): self.queue = queue self.generation_done = False def on_llm_new_token(self, token: str, **kwargs: Any) -> None: # 收到新token时放入队列 self.queue.put(token) def on_llm_end(self, *args, **kwargs: Any) -> None: # 标记LLM生成结束 self.generation_done = True # 初始化LlamaCpp LLM实例(必须开启streaming) llm = LlamaCpp( model_path="./mistral-7b-instruct-v0.1.Q4_K_M.gguf", # 替换为你的本地模型路径 temperature=0.7, max_tokens=512, top_p=0.9, streaming=True, # 关键:开启流式生成 verbose=True, ) # 定义提问模板 prompt_template = PromptTemplate( input_variables=["question"], template="你是专业助手,请回答用户问题:{question}" ) @app.post("/stream-answer") async def stream_answer(question: str): token_queue = Queue() callback = StreamingCallbackHandler(token_queue) # 为LLM绑定回调处理器 llm_with_callback = llm.bind(callbacks=[callback]) qa_chain = prompt_template | llm_with_callback # 异步生成器:从队列取token并返回给客户端 async def token_generator(): # 后台线程执行LLM调用,避免阻塞FastAPI事件循环 def run_chain(): qa_chain.invoke({"question": question}) threading.Thread(target=run_chain).start() # 持续读取队列直到生成结束且队列为空 while not callback.generation_done or not token_queue.empty(): if not token_queue.empty(): token = token_queue.get() yield token.encode("utf-8") # 转为bytes流返回 await asyncio.sleep(0.01) # 避免空轮询占用资源 return StreamingResponse(token_generator(), media_type="text/plain")
关键细节说明
StreamingCallbackHandler
- 继承LangChain的
BaseCallbackHandler,通过on_llm_new_token捕获每个生成的token,存入线程安全的Queue中。 on_llm_end标记生成完成,确保生成器能正常退出。
- 继承LangChain的
LlamaCpp配置
- 必须设置
streaming=True:这是触发逐token回调的前提,否则LLM会一次性返回完整响应。 - 确保模型路径正确,且llama-cpp-python版本≥0.2.0(支持流式回调)。
- 必须设置
FastAPI流式响应
- 使用
StreamingResponse接收异步生成器,实时返回token。 - 用后台线程执行
qa_chain.invoke:因为LangChain的LlamaCpp调用是同步的,直接在异步函数中执行会阻塞事件循环,线程包装可避免此问题。
- 使用
格式扩展(可选)
如果需要前端更易处理的SSE(Server-Sent Events)格式,可修改生成器的yield内容:yield f"data: {token}\n\n".encode("utf-8")同时将
media_type改为"text/event-stream"。
测试验证
启动FastAPI服务后,用curl测试流式输出:
curl -X POST "http://localhost:8000/stream-answer" -H "Content-Type: application/x-www-form-urlencoded" -d "question=介绍一下Mistral模型"
此时终端会逐字符打印响应,而非等待全部生成完成。
内容的提问来源于stack exchange,提问作者Maxl Gemeinderat
相关产品推荐
相关产品推荐

