如何正确返回llama_index中index.query(streaming=True)的流式查询结果?
解决Llama Index流式查询返回的两个错误
错误原因分析
第一个错误:TypeError: cannot pickle 'generator' object
调用index.query(query, streaming=True)会返回一个流式响应生成器,而非直接的结果字符串。如果直接返回这个生成器(比如在Web框架的响应流程中),框架会尝试序列化该对象,但生成器无法被pickle序列化,因此触发报错。
第二个错误:TypeError: 'StreamingResponse' object is not iterable
- 错误直接迭代
index.query(streaming=True)的返回值——该方法实际返回GeneratorStreamingResponse对象,必须调用其iter_response()方法才能获取chunk迭代器; - 错误将chunk当作字典处理,流式响应的chunk本身就是字符串,无需取
chunk["response"]; - 若使用FastAPI的
StreamingResponse,需确保传入的是有效生成器,且media_type设置合理。
正确实现代码(以FastAPI为例)
from fastapi import FastAPI, StreamingResponse from llama_index import GPTSimpleVectorIndex app = FastAPI() def stream_response(query: str, index): # 获取流式响应对象 streaming_result = index.query(query, streaming=True) # 迭代获取每个chunk并返回 for chunk in streaming_result.iter_response(): yield chunk @app.get("/stream-query") def stream_query(query: str): # 加载本地索引文件 index = GPTSimpleVectorIndex.load_from_disk("index_file.json") # 返回流式响应,media_type可根据需求改为text/html return StreamingResponse(stream_response(query, index), media_type="text/plain")
其他注意事项
- 若使用Flask等其他Web框架,流式响应的核心逻辑一致:通过
iter_response()迭代chunk,逐块写入响应; - 确保你的Llama Index版本为最新,旧版本的流式API可能存在差异。
内容的提问来源于stack exchange,提问作者Taishi Kato
相关产品推荐
相关产品推荐

