LangGraph搭配Gemini模型时工具输出无法处理(文件输入场景)
Gemini + LangGraph 工具返回结果处理异常问题(仅PDF输入时触发)
使用Gemini模型结合LangGraph框架时,LLM能成功调用工具,但无法解读或处理工具返回结果。该问题仅在PDF文件输入时出现,纯文本输入模式下运行正常。测试PDF仅包含“Apple”字样,Agent可正确识别目标公司,但无法匹配工具返回的结果。
复现代码
import asyncio import base64 import os from pathlib import Path from langchain_core.tools import tool from langchain_core.messages import HumanMessage, SystemMessage from langchain.chat_models import init_chat_model from langgraph.graph import StateGraph, START, END from langgraph.graph.message import MessagesState from langgraph.prebuilt import ToolNode # ============== CONFIGURATION ============== LLM_MODEL = "gemini-2.5-flash" LLM_MODEL_PROVIDER = "google_genai" GEMINI_API_KEY = os.environ.get("GEMINI_API_KEY") # Path to test file SCRIPT_DIR = Path(__file__).parent TEST_FILE = SCRIPT_DIR / "test.pdf" # ============== TOOLS ============== @tool def get_client(company_name: str) -> str: """ Search for a client by name. Returns matching clients or empty if not found. """ return [{"id": "f4764430-5add-446f-b35b-a9c4e272c27c", "name": "Apple"}] TOOLS = [get_client] # ============== STATE ============== class AgentState(MessagesState): """Simple state with just messages.""" pass # ============== SYSTEM PROMPT ============== SYSTEM_PROMPT = """You are an order entry assistant. Extract client information from documents. TASK: 1. Look at the provided document/text and find the client/customer name 2. Call get_client with the client name you found 3. Report the result (client name and its ID) IMPORTANT: - Read tool responses carefully - The tool returns a list of similar companies. Evaluate the results: - SUCCESS: If any result has a company name that is clearly the same entity (ignore minor formatting like spaces, periods, capitalization), that IS a match. Use that result's `id` and report success. - FAILURE: Only if the list is empty OR the names are completely different companies. """ # ============== LLM ============== llm = init_chat_model( model=LLM_MODEL, model_provider=LLM_MODEL_PROVIDER, google_api_key=GEMINI_API_KEY, ) llm_with_tools = llm.bind_tools(TOOLS) # ============== NODES ============== async def process_node(state: AgentState) -> AgentState: """Main processing node - calls LLM with tools.""" result = await llm_with_tools.ainvoke(state["messages"]) if result.tool_calls: for tc in result.tool_calls: print(f" - {tc['name']}: {tc['args']}") return {"messages": [result]} # ============== ROUTING ============== def route_tools(state: AgentState) -> str: """Route to tools if there are tool calls, otherwise to finalize.""" messages = state.get("messages", []) if not messages: return END last_message = messages[-1] if hasattr(last_message, "tool_calls") and last_message.tool_calls: print(" -> Routing to: tools") return "tools" print(" -> Routing to: finalize") return "finalize" # ============== BUILD GRAPH ============== def build_graph(): """Build the agent graph.""" builder = StateGraph(AgentState) tool_node = ToolNode(TOOLS) builder.add_node("process", process_node) builder.add_node("tools", tool_node) builder.add_edge(START, "process") builder.add_conditional_edges("process", route_tools, { "tools": "tools", "finalize": END, }) builder.add_edge("tools", "process") return builder.compile() # ============== FILE READING ============== def read_file_as_base64(filepath: Path) -> tuple[str, str]: """Read file and return (base64_content, mime_type).""" with open(filepath, "rb") as f: content = f.read() base64_content = base64.b64encode(content).decode('utf-8') # Detect mime type if content.startswith(b'%PDF'): mime_type = 'application/pdf' elif content.startswith(b'\xff\xd8\xff'): mime_type = 'image/jpeg' elif content.startswith(b'\x89PNG'): mime_type = 'image/png' else: mime_type = 'application/pdf' # default return base64_content, mime_type def create_file_message(filepath: Path) -> HumanMessage: """Create a HumanMessage with file content.""" base64_content, mime_type = read_file_as_base64(filepath) # Use proper LangChain format for multimodal content # For images: image_url type with data URL # For PDFs: Gemini supports PDFs via image_url format content = [ {"type": "text", "text": f"Process this document and find the client: {filepath.name}"}, {"type": "file", "source_type": "base64", "mime_type": mime_type, "data": base64_content} ] return HumanMessage(content=content) # ============== MAIN ============== async def run_agent(filepath: Path = None, text_input: str = None): """Run the agent with file and/or text input.""" graph = build_graph() # Build messages messages = [SystemMessage(content=SYSTEM_PROMPT)] if filepath and filepath.exists(): print(f"Reading file: {filepath}") messages.append(create_file_message(filepath)) elif text_input: messages.append(HumanMessage(content=text_input)) else: print("No input provided!") return initial_state = {"messages": messages} # Run graph config = {"configurable": {"thread_id": "debug-1"}, "recursion_limit": 20} result = await graph.ainvoke(initial_state, config=config) return result async def main(): # Try file first, fall back to text if TEST_FILE.exists(): await run_agent(filepath=TEST_FILE) else: print(f"File not found: {TEST_FILE}") print("Using text input instead...\n") await run_agent(text_input="Find the id of the client Apple") if __name__ == "__main__": asyncio.run(main())
环境依赖
langchain==1.2.0 langchain-core==1.2.4 langchain-google-genai==4.1.2 langgraph==1.0.5 google-genai==1.56.0
问题根源与修复方案
问题根源
- 多模态消息格式不兼容:当前PDF消息使用
{"type": "file"}格式,但LangChain与Gemini的多模态交互中,PDF需采用image_url类型(Gemini将PDF解析为图像序列处理)。 - 工具结果处理指令模糊:系统提示未明确要求LLM聚焦工具返回的结构化数据,PDF输入后,LLM易被多模态上下文干扰,忽略工具返回内容。
修复步骤
1. 修正PDF消息格式
修改create_file_message函数,使用Gemini兼容的image_url格式传递PDF内容:
def create_file_message(filepath: Path) -> HumanMessage: """Create a HumanMessage with file content compatible with Gemini.""" base64_content, mime_type = read_file_as_base64(filepath) # Gemini expects PDFs as image_url with data URI data_uri = f"data:{mime_type};base64,{base64_content}" content = [ {"type": "text", "text": f"Process this document and find the client: {filepath.name}"}, {"type": "image_url", "image_url": {"url": data_uri}} ] return HumanMessage(content=content)
2. 强化系统提示的工具结果处理规则
更新SYSTEM_PROMPT,明确要求LLM优先处理工具返回的结构化数据:
SYSTEM_PROMPT = """你是订单录入助手,负责从文档中提取客户信息。 任务步骤: 1. 读取提供的文档/文本,识别客户名称 2. 调用get_client工具传入识别到的客户名称 3. **重点:仔细查看工具返回的JSON格式结果,匹配客户名称后,直接输出客户名称及其ID** 关键规则: - 工具返回的是相似公司列表,需严格匹配: - 匹配成功:列表中存在名称一致(忽略大小写、空格、标点差异)的公司,直接使用其ID并报告成功 - 匹配失败:仅当列表为空或名称完全不匹配时,报告失败 - 处理工具返回时,无需再提及原始文档内容,直接基于工具结果输出"""
3. 增加调试日志(可选)
在process_node中添加工具返回结果的打印,便于排查问题:
async def process_node(state: AgentState) -> AgentState: """Main processing node - calls LLM with tools.""" result = await llm_with_tools.ainvoke(state["messages"]) if result.tool_calls: for tc in result.tool_calls: print(f" - 调用工具: {tc['name']}, 参数: {tc['args']}") # 打印工具返回的结果 if hasattr(result, 'content') and isinstance(result.content, list): for item in result.content: if item.get('type') == 'tool': print(f" - 工具返回: {item['content']}") return {"messages": [result]}
内容的提问来源于stack exchange,提问作者Gad82
相关产品推荐
相关产品推荐

