Llama3 Instruct调用工具后仅返回函数调用的问题排查求助
问题根源
你遇到的核心问题有两个:
- 每次调用
tokenizer.apply_chat_template都传入tools参数,会在对话上下文里强制插入所有工具的详细描述,强烈引导模型优先生成工具调用,完全忽略无需工具的问题。 - 缺少工具调用决策逻辑:没有判断模型输出是直接回答还是工具调用请求,导致所有场景都被当作工具调用处理。
解决方案
我们需要让模型自主判断是否需要调用工具,再根据输出分支处理:先让模型生成响应,识别是否包含工具调用标记;如果是则执行工具并生成最终回答,否则直接返回模型的回答。
修改后的完整代码
from transformers import AutoTokenizer, AutoModelForCausalLM import torch import json from typing import Dict, List # 工具函数实现 def get_current_temperature(location: str, unit: str) -> float: return 22. def get_current_wind_speed(location: str) -> float: return 6. # 工具映射:通过名称快速匹配对应函数 tool_map = { "get_current_temperature": get_current_temperature, "get_current_wind_speed": get_current_wind_speed } # 工具结构化Schema:给模型提供工具的标准描述(符合Llama 3.2规范) tools_schema = [ { "type": "function", "function": { "name": "get_current_temperature", "description": "Get the current temperature at a location.", "parameters": { "type": "object", "properties": { "location": {"type": "string", "description": "Location in 'City, Country' format"}, "unit": {"type": "string", "description": "Temperature unit", "enum": ["celsius", "fahrenheit"]} }, "required": ["location", "unit"] } } }, { "type": "function", "function": { "name": "get_current_wind_speed", "description": "Get current wind speed in km/h at a location.", "parameters": { "type": "object", "properties": { "location": {"type": "string", "description": "Location in 'City, Country' format"} }, "required": ["location"] } } } ] # 初始化模型与Tokenizer checkpoint = "models/Llama-3.2-1B-Instruct" tokenizer = AutoTokenizer.from_pretrained(checkpoint) model = AutoModelForCausalLM.from_pretrained(checkpoint, torch_dtype=torch.bfloat16, device_map="cpu") def process_conversation(messages: List[Dict]): # 第一步:让模型生成响应,自主决定是否调用工具 inputs = tokenizer.apply_chat_template( messages, tools=tools_schema, add_generation_prompt=True, return_dict=True, return_tensors="pt" ).to(model.device) outputs = model.generate( **inputs, max_new_tokens=256, temperature=0.7, do_sample=True ) response = tokenizer.decode(outputs[0][len(inputs["input_ids"][0]):], skip_special_tokens=True) # 判断是否包含工具调用标记(Llama 3.2的标准格式) if "<|tool_call_begin|>" in response and "<|tool_call_end|>" in response: # 提取并解析工具调用内容 tool_call_str = response.split("<|tool_call_begin|>")[1].split("<|tool_call_end|>")[0].strip() try: tool_call = json.loads(tool_call_str) tool_name = tool_call["name"] tool_args = tool_call["parameters"] # 执行工具函数 if tool_name in tool_map: result = tool_map[tool_name](**tool_args) # 将工具结果加入对话上下文 messages.append({"role": "assistant", "content": response}) messages.append({ "role": "tool", "name": tool_name, "content": json.dumps({"result": result}) }) # 第二步:基于工具结果生成最终回答 inputs_final = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_dict=True, return_tensors="pt" ).to(model.device) outputs_final = model.generate(**inputs_final, max_new_tokens=256) final_response = tokenizer.decode(outputs_final[0][len(inputs_final["input_ids"][0]):], skip_special_tokens=True) return final_response else: return f"未知工具:{tool_name}" except json.JSONDecodeError: return "工具调用格式解析失败" else: # 无需工具,直接返回模型的回答 return response # 测试无需工具的问题 print("测试问题:Hey, who are you?") messages = [{"role": "user", "content": "Hey, who are you?"}] print(process_conversation(messages)) # 测试需要工具的问题 print("\n测试问题:What's the temperature in Paris, France in celsius?") messages = [{"role": "user", "content": "What's the temperature in Paris, France in celsius?"}] print(process_conversation(messages))
关键修改说明
- 工具Schema标准化:用符合Llama 3.2规范的JSON结构描述工具,让模型清晰理解工具的用途和参数要求。
- 添加决策分支:通过检查模型输出中的
<|tool_call_begin|>和<|tool_call_end|>标记,区分直接回答和工具调用请求。 - 分阶段生成:第一次生成让模型判断是否调用工具,第二次生成基于工具结果输出最终回答。
- 系统提示优化(可选):可以在对话开头添加系统提示,进一步引导模型的决策逻辑:
messages = [ {"role": "system", "content": "你是一个乐于助人的助手。只有在需要时才使用提供的工具回答问题,不需要工具时直接回答。"}, {"role": "user", "content": "Hey, who are you?"} ]
内容的提问来源于stack exchange,提问作者Kodr.F
相关产品推荐
相关产品推荐

