FastAPI+Llama-index流式响应报错:'async for'需__aiter__方法
FastAPI + Llama-index 流式响应报错解决
问题场景
使用FastAPI结合Llama-index实现聊天代理的流式响应返回,后端能输出流式内容,但流结束时抛出错误:
Error: 'async for' requires an object with __aiter__ method, got NoneType
报错原因
print_response_stream()返回值无效:Llama-index的print_response_stream()仅用于打印响应流,执行完成后返回None,并非可迭代的生成器。你之前将它赋值给responseStream,迭代结束后变量变为None,导致async for触发类型错误。- 同步生成器与异步迭代不兼容:Llama-index流式响应的
response_gen是同步生成器,无法直接用async for迭代;同时同步方法chat_engine.chat会阻塞FastAPI的异步事件循环。
解决方案
修改后的完整代码
import asyncio from fastapi import FastAPI, StreamingResponse from fastapi.middleware.cors import CORSMiddleware from pydantic import BaseModel from llama_index.core import Prompt from llama_index.core.chat_engine import CondenseQuestionChatEngine # 假设index已提前初始化完成 # index = ... class InputText(BaseModel): input_text: str async def chat(input_text): custom_prompt = Prompt( """ xxx """ ) try: query_engine = index.as_query_engine(streaming=True, similarity_top_k=3) chat_engine = CondenseQuestionChatEngine.from_defaults( query_engine=query_engine, condense_question_prompt=custom_prompt, # chat_history=custom_chat_history, verbose=True, ) # 用asyncio.to_thread运行同步方法,避免阻塞异步事件循环 response = await asyncio.to_thread(chat_engine.chat, input_text) # 使用response.response_gen获取真正的token生成器,替代print_response_stream() response_gen = response.response_gen # 遍历同步生成器,逐个返回token for token in response_gen: yield token # 可选:添加微小延迟,适配前端接收节奏 await asyncio.sleep(0.01) except Exception as e: error_msg = f"Error: {str(e)}" print(error_msg) yield error_msg app = FastAPI() # 配置CORS app.add_middleware( CORSMiddleware, allow_origins=["*"], allow_methods=["GET", "POST"], allow_headers=["*"], ) @app.get("/") async def root(): return {"message": "Hello world"} @app.post("/chat") async def chat_model_001(input_text: InputText): async def generate_response(): async for token in chat(input_text.input_text): # 符合SSE标准格式,确保前端正确解析流式内容 yield f"data: {token}\n\n" return StreamingResponse(generate_response(), media_type="text/event-stream")
关键改动说明
- 替换生成器来源:舍弃
print_response_stream(),改用response.response_gen获取真正的响应token生成器。 - 异步兼容同步操作:用
asyncio.to_thread包裹chat_engine.chat,避免同步逻辑阻塞FastAPI的异步事件循环。 - 适配SSE格式:返回的每个token添加
data:前缀和换行符,符合Server-Sent Events标准,确保前端能正确接收流式响应。 - 错误处理优化:捕获异常后将错误信息返回给前端,避免流中断无提示。
内容的提问来源于stack exchange,提问作者Fran ETH
相关产品推荐
相关产品推荐

