Llama-Index多次调用query_engine.query返回空响应问题求助
问题:LlamaCPP QueryEngine二次调用返回空响应
用户代码首次执行query_engine.query()可正常获取结果,但第二次调用相同查询时,始终返回“Empty Response”。
可能的原因及解决方案
1. 模型上下文未重置
LlamaCPP实例会保留上一次生成的上下文状态,后续调用可能因上下文残留导致无法正常输出。可通过以下方式处理:
- 重置模型状态:如果使用的llama-cpp-python版本支持,调用
llm.reset()清理上下文:response = query_engine.query("How much is the Recomended Amount?") print(response) llm.reset() # 重置模型上下文 response = query_engine.query("How much is the Recomended Amount?") print(response) - 重新初始化LLM实例:对于不支持重置的场景,可在每次查询前创建新的LlamaCPP实例(注意:大模型会增加初始化开销):
def init_llm(): return LlamaCPP( model_url="https://huggingface.co/TheBloke/Llama-2-13B-chat-GGUF/resolve/main/llama-2-13b-chat.Q4_0.gguf", temperature=0.1, max_new_tokens=256, context_window=3900, generate_kwargs={}, model_kwargs={"n_gpu_layers": 1}, messages_to_prompt=messages_to_prompt, completion_to_prompt=completion_to_prompt, verbose=True, ) # 首次查询 llm = init_llm() service_context = ServiceContext.from_defaults(llm=llm, embed_model=embed_model,) index = VectorStoreIndex.from_documents(documents, service_context=service_context) query_engine = index.as_query_engine() response = query_engine.query("How much is the Recomended Amount?") print(response) # 第二次查询 llm = init_llm() service_context = ServiceContext.from_defaults(llm=llm, embed_model=embed_model,) query_engine = index.as_query_engine() response = query_engine.query("How much is the Recomended Amount?") print(response)
2. 生成参数配置缺失
generate_kwargs未设置停止词可能导致模型无法正确终止生成,返回空响应。需根据你的prompt模板添加对应停止词:
generate_kwargs={"stop": ["</s>", "Human:", "Assistant:"]}
替换为你的prompt实际使用的终止标识,确保模型能及时结束输出。
3. 上下文窗口溢出
首次查询的上下文未清理,导致第二次查询时上下文窗口占满,无法生成新内容。可尝试:
- 增大
context_window参数值,预留更多空间给新查询; - 结合
llm.reset()清理旧上下文,避免窗口溢出。
4. 依赖版本问题
旧版本的llama-cpp-python存在连续调用的bug,升级到最新稳定版:
pip install --upgrade llama-cpp-python
内容的提问来源于stack exchange,提问作者Jamie Dixon
相关产品推荐
相关产品推荐

