如何将LangChain自定义工具与Llama 2集成?代理报错及多输入代理推荐
核心问题原因
STRUCTURED_CHAT_ZERO_SHOT_REACT_DESCRIPTION代理依赖严格的JSON格式输出,而Llama 2的默认输出逻辑没有对齐OpenAI的格式要求,导致输出无法被解析器识别,进而出现"Could not parse LLM output"错误,模型也无法正确触发工具调用。
具体解决方案
1. 定制提示模板,强制输出格式
直接在系统提示中明确要求Llama 2必须遵循指定的JSON结构,示例要清晰无歧义:
- 明确区分"需要调用工具"和"直接回答"的输出格式
- 给出工具调用的JSON示例,包含所有必填参数
示例代码:
from langchain.prompts import StructuredChatPromptTemplate from langchain.agents import StructuredChatAgent, AgentExecutor # 自定义系统提示,强化格式要求 system_prompt = """ 你是具备工具调用能力的助手,必须严格按照以下规则输出: ### 工具调用格式(必须用JSON包裹) ```json {"action": "你的工具名称", "action_input": {"参数1": 值1, "参数2": 值2}}
直接回答格式
不需要调用工具时,直接输出最终答案即可。
可用工具列表:
{tools}
工具详细说明:
{tool_names}: {tool_descriptions}
"""
prompt = StructuredChatPromptTemplate.from_messages(
[
("system", system_prompt),
("human", "{input}"),
("ai", "{agent_scratchpad}"),
]
)
agent = StructuredChatAgent.from_llm_and_tools(
llm=llama2_llm, # 你的Llama 2 LLM实例
tools=calculator_tools, # 你的计算器工具
prompt=prompt
)
agent_executor = AgentExecutor(agent=agent, tools=calculator_tools, verbose=True)
### 2. 调整LLM参数,降低输出随机性 Llama 2的高温度参数会导致输出不稳定,调整以下参数提升格式一致性: - 设置`temperature=0`:强制模型生成确定性输出,减少格式偏差 - 添加`stop`参数:指定停止符(如`["\nObservation:"]`),避免输出多余内容干扰解析 示例代码: ```python from langchain.llms import HuggingFacePipeline # 初始化Llama 2时调整参数 llama2_llm = HuggingFacePipeline.from_model_id( model_id="meta-llama/Llama-2-7b-chat-hf", task="text-generation", pipeline_kwargs={ "temperature": 0, "max_new_tokens": 512, "stop": ["\nObservation:", "\nHuman:"] } )
3. 自定义输出解析器,兼容Llama 2的输出
Llama 2可能输出多余的前缀(如"答:")或代码块标记,自定义解析器清理后再解析:
from langchain.agents import AgentOutputParser from langchain.schema import AgentAction, AgentFinish import json class Llama2StructuredParser(AgentOutputParser): def parse(self, llm_output: str) -> AgentAction | AgentFinish: # 清理输出中的无关内容 cleaned_output = llm_output.strip() cleaned_output = cleaned_output.replace("```json", "").replace("```", "") try: output_data = json.loads(cleaned_output) # 识别工具调用 if "action" in output_data and "action_input" in output_data: return AgentAction( tool=output_data["action"], tool_input=output_data["action_input"], log=llm_output ) # 直接回答 else: return AgentFinish( return_values={"output": cleaned_output}, log=llm_output ) except json.JSONDecodeError: # 解析失败时默认直接返回内容 return AgentFinish( return_values={"output": llm_output}, log=llm_output ) # 给代理绑定自定义解析器 agent = StructuredChatAgent.from_llm_and_tools( llm=llama2_llm, tools=calculator_tools, prompt=prompt, output_parser=Llama2StructuredParser() )
多输入场景的替代代理推荐
如果STRUCTURED_CHAT代理适配成本过高,推荐以下更适合多输入场景的代理:
1. FUNCTIONS代理
适合支持函数调用的Llama 2微调版本(如Llama-2-7b-chat-hf),只需将工具定义为函数格式,模型可自动生成包含多参数的调用请求,对齐函数调用规范。
2. ZERO_SHOT_REACT_DESCRIPTION代理
无需严格JSON格式,依赖ReAct思维链格式(Thought/Action/Action Input/Observation),Llama 2更容易遵循这种自然语言格式,适配成本低,支持多输入参数的工具调用。
3. CONVERSATIONAL_REACT_DESCRIPTION代理
针对多轮对话场景设计,可保留上下文信息,同时支持多输入工具调用,适合需要连续交互的场景。
内容的提问来源于stack exchange,提问作者Prabhjot Kaur

