You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何正确返回llama_index中index.query(streaming=True)的流式查询结果?

解决Llama Index流式查询返回的两个错误

错误原因分析

第一个错误:TypeError: cannot pickle 'generator' object

调用index.query(query, streaming=True)会返回一个流式响应生成器,而非直接的结果字符串。如果直接返回这个生成器(比如在Web框架的响应流程中),框架会尝试序列化该对象,但生成器无法被pickle序列化,因此触发报错。

第二个错误:TypeError: 'StreamingResponse' object is not iterable

  1. 错误直接迭代index.query(streaming=True)的返回值——该方法实际返回GeneratorStreamingResponse对象,必须调用其iter_response()方法才能获取chunk迭代器;
  2. 错误将chunk当作字典处理,流式响应的chunk本身就是字符串,无需取chunk["response"];
  3. 若使用FastAPI的StreamingResponse,需确保传入的是有效生成器,且media_type设置合理。

正确实现代码(以FastAPI为例)

from fastapi import FastAPI, StreamingResponse
from llama_index import GPTSimpleVectorIndex

app = FastAPI()

def stream_response(query: str, index):
    # 获取流式响应对象
    streaming_result = index.query(query, streaming=True)
    # 迭代获取每个chunk并返回
    for chunk in streaming_result.iter_response():
        yield chunk

@app.get("/stream-query")
def stream_query(query: str):
    # 加载本地索引文件
    index = GPTSimpleVectorIndex.load_from_disk("index_file.json")
    # 返回流式响应,media_type可根据需求改为text/html
    return StreamingResponse(stream_response(query, index), media_type="text/plain")

其他注意事项

  • 若使用Flask等其他Web框架,流式响应的核心逻辑一致:通过iter_response()迭代chunk,逐块写入响应;
  • 确保你的Llama Index版本为最新,旧版本的流式API可能存在差异。

内容的提问来源于stack exchange,提问作者Taishi Kato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 12:58:14