You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在使用LangGraph/LangChain时,如何让工具返回图片?

在使用LangGraph/LangChain时,如何让工具返回图片?

嘿,我来给你捋捋怎么实现这个需求。核心思路是让工具输出的内容能被多模态LLM正确识别为图片,同时确保LLM能获取到图片的有效数据。下面分两种常用方案来讲解:


方案一:返回Base64编码的图片(推荐)

这种方式不需要LLM访问本地文件系统,兼容性更强,适合大多数多模态LLM(比如GPT-4V、Claude 3等)。

调整工具代码

我们可以把截图转换成Base64字符串,再用结构化的格式返回,让LLM一眼就能识别这是图片数据:

import base64
from langchain_core.tools import tool

def capture_screenshot_to_base64():
    # 这里替换成你实际的截图逻辑,生成图片文件
    img_path = "/path/to/pic.png"
    
    # 将图片编码为Base64字符串
    with open(img_path, "rb") as img_file:
        base64_str = base64.b64encode(img_file.read()).decode("utf-8")
    
    # 返回结构化结果,明确标记类型和格式
    return {
        "type": "image",
        "format": "png",
        "data": base64_str,
        "prompt": "请分析这张网页截图的内容"
    }

@tool
def get_screenshot():
    """获取当前网页的截图,返回Base64编码格式的图片数据"""
    return capture_screenshot_to_base64()

在LangGraph中处理结果

接下来要把工具返回的Base64数据转换成LLM能理解的消息格式。比如在LLM节点里构造包含图片的对话消息:

from langchain_core.messages import HumanMessage
from langgraph.graph import StateGraph

def llm_process_node(state):
    # 取出工具返回的最新结果
    tool_output = state["messages"][-1].content
    
    # 构造多模态消息
    if isinstance(tool_output, dict) and tool_output["type"] == "image":
        message = HumanMessage(
            content=[
                {"type": "text", "text": tool_output["prompt"]},
                {"type": "image_url", "image_url": {
                    "url": f"data:image/{tool_output['format']};base64,{tool_output['data']}"
                }}
            ]
        )
    else:
        message = HumanMessage(content=str(tool_output))
    
    # 调用多模态LLM处理
    response = your_multimodal_llm.invoke([message])
    return {"messages": [response]}

# 后续可以继续构建LangGraph的节点和边逻辑...

方案二:返回本地图片路径(仅限LLM能访问文件系统的场景)

如果你的LLM部署在本地,或者能直接访问存储图片的路径,也可以直接返回文件路径,但需要明确提示LLM去读取:

from langchain_core.tools import tool

def capture_screenshot_to_file():
    # 截图并保存到本地路径
    return "/path/to/pic.png"

@tool
def get_screenshot():
    """获取当前网页的截图,返回本地文件路径"""
    img_path = capture_screenshot_to_file()
    return f"网页截图已保存到路径:{img_path},请读取该文件并分析截图内容"

这种方式的缺点是依赖LLM的环境权限,比如云部署的LLM可能无法访问你的本地文件,所以更适合本地开发测试场景。


备注:内容来源于stack exchange,提问作者Letian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 17:14:30