You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LangGraph搭配Gemini模型时工具输出无法处理(文件输入场景)

Gemini + LangGraph 工具返回结果处理异常问题(仅PDF输入时触发)

使用Gemini模型结合LangGraph框架时,LLM能成功调用工具,但无法解读或处理工具返回结果。该问题仅在PDF文件输入时出现,纯文本输入模式下运行正常。测试PDF仅包含“Apple”字样,Agent可正确识别目标公司,但无法匹配工具返回的结果。

复现代码

import asyncio
import base64
import os
from pathlib import Path
from langchain_core.tools import tool
from langchain_core.messages import HumanMessage, SystemMessage
from langchain.chat_models import init_chat_model
from langgraph.graph import StateGraph, START, END
from langgraph.graph.message import MessagesState
from langgraph.prebuilt import ToolNode

# ============== CONFIGURATION ==============
LLM_MODEL = "gemini-2.5-flash"
LLM_MODEL_PROVIDER = "google_genai"
GEMINI_API_KEY = os.environ.get("GEMINI_API_KEY")

# Path to test file
SCRIPT_DIR = Path(__file__).parent
TEST_FILE = SCRIPT_DIR / "test.pdf"

# ============== TOOLS ==============
@tool
def get_client(company_name: str) -> str:
    """
    Search for a client by name.
    Returns matching clients or empty if not found.
    """
    return [{"id": "f4764430-5add-446f-b35b-a9c4e272c27c", "name": "Apple"}]


TOOLS = [get_client]


# ============== STATE ==============
class AgentState(MessagesState):
    """Simple state with just messages."""
    pass


# ============== SYSTEM PROMPT ==============
SYSTEM_PROMPT = """You are an order entry assistant. Extract client information from documents.

TASK:
1. Look at the provided document/text and find the client/customer name
2. Call get_client with the client name you found
3. Report the result (client name and its ID)

IMPORTANT: 
- Read tool responses carefully
- The tool returns a list of similar companies. Evaluate the results:
    - SUCCESS: If any result has a company name that is clearly the same entity (ignore minor formatting like spaces, periods, capitalization), that IS a match. Use that result's `id` and report success.
    - FAILURE: Only if the list is empty OR the names are completely different companies.
"""


# ============== LLM ==============
llm = init_chat_model(
    model=LLM_MODEL,
    model_provider=LLM_MODEL_PROVIDER,
    google_api_key=GEMINI_API_KEY,
)
llm_with_tools = llm.bind_tools(TOOLS)

# ============== NODES ==============
async def process_node(state: AgentState) -> AgentState:
    """Main processing node - calls LLM with tools."""    
    result = await llm_with_tools.ainvoke(state["messages"])
    
    if result.tool_calls:
        for tc in result.tool_calls:
            print(f"      - {tc['name']}: {tc['args']}")
    
    return {"messages": [result]}

# ============== ROUTING ==============
def route_tools(state: AgentState) -> str:
    """Route to tools if there are tool calls, otherwise to finalize."""
    messages = state.get("messages", [])
    if not messages:
        return END
    
    last_message = messages[-1]
    if hasattr(last_message, "tool_calls") and last_message.tool_calls:
        print("    -> Routing to: tools")
        return "tools"
    
    print("    -> Routing to: finalize")
    return "finalize"

# ============== BUILD GRAPH ==============
def build_graph():
    """Build the agent graph."""
    builder = StateGraph(AgentState)
    
    tool_node = ToolNode(TOOLS)
    builder.add_node("process", process_node)
    builder.add_node("tools", tool_node)
    builder.add_edge(START, "process")
    builder.add_conditional_edges("process", route_tools, {
        "tools": "tools",
        "finalize": END,
    })
    builder.add_edge("tools", "process")
    
    return builder.compile()

# ============== FILE READING ==============
def read_file_as_base64(filepath: Path) -> tuple[str, str]:
    """Read file and return (base64_content, mime_type)."""
    with open(filepath, "rb") as f:
        content = f.read()
    
    base64_content = base64.b64encode(content).decode('utf-8')
    
    # Detect mime type
    if content.startswith(b'%PDF'):
        mime_type = 'application/pdf'
    elif content.startswith(b'\xff\xd8\xff'):
        mime_type = 'image/jpeg'
    elif content.startswith(b'\x89PNG'):
        mime_type = 'image/png'
    else:
        mime_type = 'application/pdf'  # default
    
    return base64_content, mime_type

def create_file_message(filepath: Path) -> HumanMessage:
    """Create a HumanMessage with file content."""
    base64_content, mime_type = read_file_as_base64(filepath)
    
    # Use proper LangChain format for multimodal content
    # For images: image_url type with data URL
    # For PDFs: Gemini supports PDFs via image_url format
    content = [
        {"type": "text", "text": f"Process this document and find the client: {filepath.name}"},
        {"type": "file", "source_type": "base64", "mime_type": mime_type, "data": base64_content}
    ]
    
    return HumanMessage(content=content)

# ============== MAIN ==============
async def run_agent(filepath: Path = None, text_input: str = None):
    """Run the agent with file and/or text input."""
    graph = build_graph()
    
    # Build messages
    messages = [SystemMessage(content=SYSTEM_PROMPT)]
    
    if filepath and filepath.exists():
        print(f"Reading file: {filepath}")
        messages.append(create_file_message(filepath))
    elif text_input:
        messages.append(HumanMessage(content=text_input))
    else:
        print("No input provided!")
        return
    
    initial_state = {"messages": messages}
    
    # Run graph
    config = {"configurable": {"thread_id": "debug-1"}, "recursion_limit": 20}
    result = await graph.ainvoke(initial_state, config=config)   
    return result


async def main():
    # Try file first, fall back to text
    if TEST_FILE.exists():
        await run_agent(filepath=TEST_FILE)
    else:
        print(f"File not found: {TEST_FILE}")
        print("Using text input instead...\n")
        await run_agent(text_input="Find the id of the client Apple")


if __name__ == "__main__":
    asyncio.run(main())

环境依赖

langchain==1.2.0
langchain-core==1.2.4
langchain-google-genai==4.1.2
langgraph==1.0.5
google-genai==1.56.0

问题根源与修复方案

问题根源

  1. 多模态消息格式不兼容:当前PDF消息使用{"type": "file"}格式,但LangChain与Gemini的多模态交互中,PDF需采用image_url类型(Gemini将PDF解析为图像序列处理)。
  2. 工具结果处理指令模糊:系统提示未明确要求LLM聚焦工具返回的结构化数据,PDF输入后,LLM易被多模态上下文干扰,忽略工具返回内容。

修复步骤

1. 修正PDF消息格式

修改create_file_message函数,使用Gemini兼容的image_url格式传递PDF内容:

def create_file_message(filepath: Path) -> HumanMessage:
    """Create a HumanMessage with file content compatible with Gemini."""
    base64_content, mime_type = read_file_as_base64(filepath)
    
    # Gemini expects PDFs as image_url with data URI
    data_uri = f"data:{mime_type};base64,{base64_content}"
    content = [
        {"type": "text", "text": f"Process this document and find the client: {filepath.name}"},
        {"type": "image_url", "image_url": {"url": data_uri}}
    ]
    
    return HumanMessage(content=content)

2. 强化系统提示的工具结果处理规则

更新SYSTEM_PROMPT,明确要求LLM优先处理工具返回的结构化数据:

SYSTEM_PROMPT = """你是订单录入助手,负责从文档中提取客户信息。

任务步骤:
1. 读取提供的文档/文本,识别客户名称
2. 调用get_client工具传入识别到的客户名称
3. **重点:仔细查看工具返回的JSON格式结果,匹配客户名称后,直接输出客户名称及其ID**

关键规则:
- 工具返回的是相似公司列表,需严格匹配:
  - 匹配成功:列表中存在名称一致(忽略大小写、空格、标点差异)的公司,直接使用其ID并报告成功
  - 匹配失败:仅当列表为空或名称完全不匹配时,报告失败
- 处理工具返回时,无需再提及原始文档内容,直接基于工具结果输出"""

3. 增加调试日志(可选)

在process_node中添加工具返回结果的打印,便于排查问题:

async def process_node(state: AgentState) -> AgentState:
    """Main processing node - calls LLM with tools."""    
    result = await llm_with_tools.ainvoke(state["messages"])
    
    if result.tool_calls:
        for tc in result.tool_calls:
            print(f"      - 调用工具: {tc['name']}, 参数: {tc['args']}")
    # 打印工具返回的结果
    if hasattr(result, 'content') and isinstance(result.content, list):
        for item in result.content:
            if item.get('type') == 'tool':
                print(f"      - 工具返回: {item['content']}")
    
    return {"messages": [result]}

内容的提问来源于stack exchange,提问作者Gad82

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.11 19:14:49