Haystack 2.12非Agent RAG Pipeline添加DuckDuckGo搜索工具失败求助
问题分析与修复方案
核心问题拆解
- LLM无工具调用意识:当前Prompt完全未告知模型存在WebSearch工具,导致模型无法判断何时需要触发搜索,只会返回预设的无实时数据提示。
- 工具调用流程断裂:即便触发工具调用,搜索结果也未回传给LLM,无法基于新数据生成回答。
- LLM未配置工具支持:Haystack 2.x中,OpenAIGenerator需显式传入
tools参数,才能让模型生成符合格式的工具调用指令。 - Pipeline路由不完整:缺少工具调用结果到LLM的回流路径,也未明确最终回复的输出端口。
具体修复步骤
1. 修改Prompt模板,添加工具调用规则
在Prompt中明确告知模型工具的存在、适用场景和调用格式:
def _create_system_prompt(self): return """ You are an AI assistant designed to provide **clear, concise, and accurate** responses. Your goal is to **help users efficiently** while providing recommendations only if the question relates to something you can recommend. Your responses should be **direct, informative, and polite**, without unnecessary details or filler content. **可用工具**: 当你需要实时数据或当前文档中没有的信息时,必须调用WebSearch工具,调用格式严格遵循: <|tool_call_begin|>[{"name": "WebSearch", "parameters": {"query": "你的搜索关键词"}}]<|tool_call_end|> **Conversation Context**: {% if conversation_history %} Previous conversation history: {{ conversation_history }} {% else %} This is a new conversation. {% endif %} **Documents**: {% for doc in documents %} {{ doc.content }} {% endfor %} **Image Description** {% if image_description %} An image description has been provided. Use the description to assist with the response, including any user-related details that might be inferred. {{ image_description }} {% else %} No image description provided. {% endif %} **Question**: {{ question }} **Answer**: """
2. 配置OpenAIGenerator支持工具调用
初始化LLM时传入tools参数,让模型知晓可用工具:
def _initialize_language_model(self): api_key = "API_KEY" model_name = self.model # 传入WebTool,启用模型的工具调用能力 return OpenAIGenerator(api_key=Secret.from_token(api_key), model=model_name, tools=[self.web_tool])
3. 完善Pipeline的路由与连接
补充工具调用结果到LLM的回流路径,并设置输出端口:
def _initialize_rag_pipeline(self): pipeline = Pipeline() pipeline.add_component("retriever", self.retriever) pipeline.add_component("prompt_builder", self.prompt_builder) pipeline.add_component("llm", self.llm) pipeline.add_component("router", ConditionalRouter(self.web_route)) pipeline.add_component("tool_invoker", ToolInvoker(tools=[self.web_tool])) # 原有连接保留 pipeline.connect("retriever.documents", "prompt_builder.documents") pipeline.connect("prompt_builder", "llm.prompt") pipeline.connect("llm.replies", "router.replies") pipeline.connect("router.there_are_tool_calls", "tool_invoker.messages") # 新增:工具调用结果回传给LLM,生成最终回答 pipeline.connect("tool_invoker.replies", "llm.messages") # 设置输出端口,覆盖两种回复场景 pipeline.set_outputs(["router.final_replies", "llm.replies"]) return pipeline
4. 添加Pipeline执行与结果处理方法
新增run方法处理输入输出,判断返回直接回复或工具调用后的结果:
def run(self, question: str, conversation_history: List[ChatMessage] = None, image_description: str = None): inputs = { "retriever": {"query": question}, "prompt_builder": { "question": question, "conversation_history": conversation_history or [], "image_description": image_description or "" } } result = self.rag_pipeline.run(inputs) # 优先返回无需工具调用的直接回复 if "final_replies" in result and result["final_replies"]: return result["final_replies"][0].content # 其次返回工具调用后的生成回复 elif "replies" in result and result["replies"]: return result["replies"][0].content # 兜底返回无数据提示 else: return "I didn't have real-time data"
5. 调整类初始化顺序
由于LLM需要依赖WebTool,需提前初始化web_tool:
def __init__(self, model="gpt-4o-mini"): self.model = model self.prompt_template = self._create_system_prompt() self.document_store = self._initialize_document_store() self.web_tool = self._initialize_web_tool() # 提前初始化,供LLM使用 self.llm = self._initialize_language_model() self.retriever = self._initialize_retriever() self.prompt_builder = self._initialize_prompt_builder() self.web_route = self._initialize_web_route() self.rag_pipeline = self._initialize_rag_pipeline()
内容的提问来源于stack exchange,提问作者Abstract
相关产品推荐
相关产品推荐

