You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastAPI+Llama-index流式响应报错:'async for'需__aiter__方法

FastAPI + Llama-index 流式响应报错解决

问题场景

使用FastAPI结合Llama-index实现聊天代理的流式响应返回,后端能输出流式内容,但流结束时抛出错误:

Error: 'async for' requires an object with __aiter__ method, got NoneType

报错原因

  1. print_response_stream()返回值无效:Llama-index的print_response_stream()仅用于打印响应流,执行完成后返回None,并非可迭代的生成器。你之前将它赋值给responseStream,迭代结束后变量变为None,导致async for触发类型错误。
  2. 同步生成器与异步迭代不兼容:Llama-index流式响应的response_gen是同步生成器,无法直接用async for迭代;同时同步方法chat_engine.chat会阻塞FastAPI的异步事件循环。

解决方案

修改后的完整代码

import asyncio
from fastapi import FastAPI, StreamingResponse
from fastapi.middleware.cors import CORSMiddleware
from pydantic import BaseModel
from llama_index.core import Prompt
from llama_index.core.chat_engine import CondenseQuestionChatEngine

# 假设index已提前初始化完成
# index = ...

class InputText(BaseModel):
    input_text: str

async def chat(input_text):
    custom_prompt = Prompt(
        """ xxx """
    )

    try:
        query_engine = index.as_query_engine(streaming=True, similarity_top_k=3)
        chat_engine = CondenseQuestionChatEngine.from_defaults(
            query_engine=query_engine,
            condense_question_prompt=custom_prompt,
            # chat_history=custom_chat_history,
            verbose=True,
        )
        # 用asyncio.to_thread运行同步方法,避免阻塞异步事件循环
        response = await asyncio.to_thread(chat_engine.chat, input_text)
        # 使用response.response_gen获取真正的token生成器,替代print_response_stream()
        response_gen = response.response_gen
        
        # 遍历同步生成器,逐个返回token
        for token in response_gen:
            yield token
            # 可选:添加微小延迟,适配前端接收节奏
            await asyncio.sleep(0.01)
    except Exception as e:
        error_msg = f"Error: {str(e)}"
        print(error_msg)
        yield error_msg


app = FastAPI()

# 配置CORS
app.add_middleware(
    CORSMiddleware,
    allow_origins=["*"],
    allow_methods=["GET", "POST"],
    allow_headers=["*"],
)


@app.get("/")
async def root():
    return {"message": "Hello world"}


@app.post("/chat")
async def chat_model_001(input_text: InputText):
    async def generate_response():
        async for token in chat(input_text.input_text):
            # 符合SSE标准格式,确保前端正确解析流式内容
            yield f"data: {token}\n\n"

    return StreamingResponse(generate_response(), media_type="text/event-stream")

关键改动说明

  1. 替换生成器来源:舍弃print_response_stream(),改用response.response_gen获取真正的响应token生成器。
  2. 异步兼容同步操作:用asyncio.to_thread包裹chat_engine.chat,避免同步逻辑阻塞FastAPI的异步事件循环。
  3. 适配SSE格式:返回的每个token添加data: 前缀和换行符,符合Server-Sent Events标准,确保前端能正确接收流式响应。
  4. 错误处理优化:捕获异常后将错误信息返回给前端,避免流中断无提示。

内容的提问来源于stack exchange,提问作者Fran ETH

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 00:53:31