You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Llama3 Instruct调用工具后仅返回函数调用的问题排查求助

问题根源

你遇到的核心问题有两个:

  1. 每次调用tokenizer.apply_chat_template都传入tools参数,会在对话上下文里强制插入所有工具的详细描述,强烈引导模型优先生成工具调用,完全忽略无需工具的问题。
  2. 缺少工具调用决策逻辑:没有判断模型输出是直接回答还是工具调用请求,导致所有场景都被当作工具调用处理。
解决方案

我们需要让模型自主判断是否需要调用工具,再根据输出分支处理:先让模型生成响应,识别是否包含工具调用标记;如果是则执行工具并生成最终回答,否则直接返回模型的回答。

修改后的完整代码

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
import json
from typing import Dict, List

# 工具函数实现
def get_current_temperature(location: str, unit: str) -> float:
    return 22.

def get_current_wind_speed(location: str) -> float:
    return 6.

# 工具映射:通过名称快速匹配对应函数
tool_map = {
    "get_current_temperature": get_current_temperature,
    "get_current_wind_speed": get_current_wind_speed
}

# 工具结构化Schema:给模型提供工具的标准描述(符合Llama 3.2规范)
tools_schema = [
    {
        "type": "function",
        "function": {
            "name": "get_current_temperature",
            "description": "Get the current temperature at a location.",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Location in 'City, Country' format"},
                    "unit": {"type": "string", "description": "Temperature unit", "enum": ["celsius", "fahrenheit"]}
                },
                "required": ["location", "unit"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "get_current_wind_speed",
            "description": "Get current wind speed in km/h at a location.",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Location in 'City, Country' format"}
                },
                "required": ["location"]
            }
        }
    }
]

# 初始化模型与Tokenizer
checkpoint = "models/Llama-3.2-1B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
model = AutoModelForCausalLM.from_pretrained(checkpoint, torch_dtype=torch.bfloat16, device_map="cpu")

def process_conversation(messages: List[Dict]):
    # 第一步:让模型生成响应,自主决定是否调用工具
    inputs = tokenizer.apply_chat_template(
        messages,
        tools=tools_schema,
        add_generation_prompt=True,
        return_dict=True,
        return_tensors="pt"
    ).to(model.device)
    
    outputs = model.generate(
        **inputs,
        max_new_tokens=256,
        temperature=0.7,
        do_sample=True
    )
    
    response = tokenizer.decode(outputs[0][len(inputs["input_ids"][0]):], skip_special_tokens=True)
    
    # 判断是否包含工具调用标记(Llama 3.2的标准格式)
    if "<|tool_call_begin|>" in response and "<|tool_call_end|>" in response:
        # 提取并解析工具调用内容
        tool_call_str = response.split("<|tool_call_begin|>")[1].split("<|tool_call_end|>")[0].strip()
        try:
            tool_call = json.loads(tool_call_str)
            tool_name = tool_call["name"]
            tool_args = tool_call["parameters"]
            
            # 执行工具函数
            if tool_name in tool_map:
                result = tool_map[tool_name](**tool_args)
                # 将工具结果加入对话上下文
                messages.append({"role": "assistant", "content": response})
                messages.append({
                    "role": "tool",
                    "name": tool_name,
                    "content": json.dumps({"result": result})
                })
                
                # 第二步:基于工具结果生成最终回答
                inputs_final = tokenizer.apply_chat_template(
                    messages,
                    add_generation_prompt=True,
                    return_dict=True,
                    return_tensors="pt"
                ).to(model.device)
                
                outputs_final = model.generate(**inputs_final, max_new_tokens=256)
                final_response = tokenizer.decode(outputs_final[0][len(inputs_final["input_ids"][0]):], skip_special_tokens=True)
                return final_response
            else:
                return f"未知工具:{tool_name}"
        except json.JSONDecodeError:
            return "工具调用格式解析失败"
    else:
        # 无需工具,直接返回模型的回答
        return response

# 测试无需工具的问题
print("测试问题:Hey, who are you?")
messages = [{"role": "user", "content": "Hey, who are you?"}]
print(process_conversation(messages))

# 测试需要工具的问题
print("\n测试问题:What's the temperature in Paris, France in celsius?")
messages = [{"role": "user", "content": "What's the temperature in Paris, France in celsius?"}]
print(process_conversation(messages))

关键修改说明

  1. 工具Schema标准化:用符合Llama 3.2规范的JSON结构描述工具,让模型清晰理解工具的用途和参数要求。
  2. 添加决策分支:通过检查模型输出中的<|tool_call_begin|>和<|tool_call_end|>标记,区分直接回答和工具调用请求。
  3. 分阶段生成:第一次生成让模型判断是否调用工具,第二次生成基于工具结果输出最终回答。
  4. 系统提示优化(可选):可以在对话开头添加系统提示,进一步引导模型的决策逻辑:
    messages = [
        {"role": "system", "content": "你是一个乐于助人的助手。只有在需要时才使用提供的工具回答问题,不需要工具时直接回答。"},
        {"role": "user", "content": "Hey, who are you?"}
    ]
    

内容的提问来源于stack exchange,提问作者Kodr.F

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:33:19