在使用LangGraph/LangChain时,如何让工具返回图片?
在使用LangGraph/LangChain时,如何让工具返回图片?
嘿,我来给你捋捋怎么实现这个需求。核心思路是让工具输出的内容能被多模态LLM正确识别为图片,同时确保LLM能获取到图片的有效数据。下面分两种常用方案来讲解:
方案一:返回Base64编码的图片(推荐)
这种方式不需要LLM访问本地文件系统,兼容性更强,适合大多数多模态LLM(比如GPT-4V、Claude 3等)。
调整工具代码
我们可以把截图转换成Base64字符串,再用结构化的格式返回,让LLM一眼就能识别这是图片数据:
import base64 from langchain_core.tools import tool def capture_screenshot_to_base64(): # 这里替换成你实际的截图逻辑,生成图片文件 img_path = "/path/to/pic.png" # 将图片编码为Base64字符串 with open(img_path, "rb") as img_file: base64_str = base64.b64encode(img_file.read()).decode("utf-8") # 返回结构化结果,明确标记类型和格式 return { "type": "image", "format": "png", "data": base64_str, "prompt": "请分析这张网页截图的内容" } @tool def get_screenshot(): """获取当前网页的截图,返回Base64编码格式的图片数据""" return capture_screenshot_to_base64()
在LangGraph中处理结果
接下来要把工具返回的Base64数据转换成LLM能理解的消息格式。比如在LLM节点里构造包含图片的对话消息:
from langchain_core.messages import HumanMessage from langgraph.graph import StateGraph def llm_process_node(state): # 取出工具返回的最新结果 tool_output = state["messages"][-1].content # 构造多模态消息 if isinstance(tool_output, dict) and tool_output["type"] == "image": message = HumanMessage( content=[ {"type": "text", "text": tool_output["prompt"]}, {"type": "image_url", "image_url": { "url": f"data:image/{tool_output['format']};base64,{tool_output['data']}" }} ] ) else: message = HumanMessage(content=str(tool_output)) # 调用多模态LLM处理 response = your_multimodal_llm.invoke([message]) return {"messages": [response]} # 后续可以继续构建LangGraph的节点和边逻辑...
方案二:返回本地图片路径(仅限LLM能访问文件系统的场景)
如果你的LLM部署在本地,或者能直接访问存储图片的路径,也可以直接返回文件路径,但需要明确提示LLM去读取:
from langchain_core.tools import tool def capture_screenshot_to_file(): # 截图并保存到本地路径 return "/path/to/pic.png" @tool def get_screenshot(): """获取当前网页的截图,返回本地文件路径""" img_path = capture_screenshot_to_file() return f"网页截图已保存到路径:{img_path},请读取该文件并分析截图内容"
这种方式的缺点是依赖LLM的环境权限,比如云部署的LLM可能无法访问你的本地文件,所以更适合本地开发测试场景。
备注:内容来源于stack exchange,提问作者Letian
相关产品推荐
相关产品推荐

