如何通过LlamaIndex向FastAPI端点发送流式响应?
在FastAPI中实现LlamaIndex流式响应
要让FastAPI端点返回LlamaIndex的流式响应,需要结合FastAPI的StreamingResponse和异步处理逻辑,以下是具体实现方案:
核心实现步骤
- 导入FastAPI的
StreamingResponse,用于包装流式生成器 - 将LlamaIndex的同步
stream_chat方法放到线程中执行,避免阻塞FastAPI的事件循环 - 定义生成器函数,逐个返回LlamaIndex输出的token
完整代码示例
from fastapi import APIRouter from fastapi.responses import StreamingResponse import asyncio # 确保已导入你的QuestionInput模型和cache对象 bot_router = APIRouter() @bot_router.post("/bot/pdf_convo") async def pdf_convo(query: QuestionInput): chat_engine = cache["chat_engine"] user_question = query.content # 异步生成器函数,流式返回响应token async def stream_tokens(): # 用asyncio.to_thread包装同步的stream_chat方法,避免阻塞事件循环 streaming_response = await asyncio.to_thread(chat_engine.stream_chat, user_question) for token in streaming_response.response_gen: # 直接返回token,若需SSE格式可改为 yield f"data: {token}\n\n" yield token # 返回StreamingResponse,指定媒体类型 return StreamingResponse(stream_tokens(), media_type="text/plain")
关键细节说明
- 异步处理:LlamaIndex的
stream_chat是同步方法,必须用asyncio.to_thread包装后才能在异步路由中安全使用,否则会阻塞FastAPI的事件循环 - 媒体类型适配:如果前端需要Server-Sent Events(SSE)格式,将
media_type改为"text/event-stream",同时调整yield的内容格式为SSE规范(如yield f"data: {token}\n\n") - 前端对接:前端需要使用支持流式接收的方式请求接口,比如JavaScript的
fetch结合ReadableStream来逐段处理返回的token
内容的提问来源于stack exchange,提问作者Mubashir Ahmed Siddiqui
相关产品推荐
相关产品推荐

