You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

FastMCP图像处理方案及Pillow搭建MCP服务器的图像表示格式困惑

FastMCP图像处理方案及Pillow搭建MCP服务器的图像表示格式困惑

Hey there! I totally get where you're coming from—this is a super common pain point when building MCP servers with Pillow and working alongside LLMs. Let's break down some practical, actionable solutions that balance data accessibility for your server and processing efficiency for the LLM:

1. 轻量级图像特征摘要 + 本地ID-路径映射表

This is my go-to for most cases where the LLM doesn't need full visual detail, just core image properties to make decisions.

  • 核心思路:用Pillow提取图像的关键属性生成简洁的文本摘要,再给每个图像分配唯一ID,在服务器本地维护ID与实际文件路径的映射表。LLM只接收轻量化的摘要,MCP服务器则通过ID快速关联到原始图像文件。
  • 代码示例:
    from PIL import Image
    import numpy as np
    import json
    import uuid
    
    # 本地缓存:用于映射图像ID和文件路径
    image_cache = {}
    
    def process_and_register_image(img_path):
        # 生成短版唯一ID
        img_id = str(uuid.uuid4())[:8]
        # 用Pillow提取图像核心特征
        with Image.open(img_path) as img:
            width, height = img.size
            mode = img.mode
            # 提取中心区域平均RGB作为主色调简化版
            crop_size = min(10, width, height)
            center_x, center_y = width//2, height//2
            crop = img.crop((
                center_x - crop_size//2,
                center_y - crop_size//2,
                center_x + crop_size//2,
                center_y + crop_size//2
            ))
            avg_rgb = tuple(np.mean(np.array(crop), axis=(0,1)).astype(int))
        
        # 生成给LLM的摘要文本
        summary = f"Image ID: {img_id} | 尺寸: {width}x{height} | 模式: {mode} | 主色调RGB: {avg_rgb}"
        # 注册到本地缓存
        image_cache[img_id] = img_path
        # 持久化缓存到文件(可选)
        with open("image_cache.json", "w") as f:
            json.dump(image_cache, f)
        return summary
    
    # 使用示例
    llm_input = process_and_register_image("./my_target_image.png")
    
  • 优势:传给LLM的文本极短,完全避免Base64的臃肿问题,同时服务器能随时通过ID获取完整图像数据,兼顾效率与实用性。

2. 压缩缩略图的Base64格式

如果LLM需要基础视觉上下文(比如区分是照片还是示意图),这个方案能在数据量和视觉信息量之间取得平衡。

  • 核心思路:先用Pillow把图像压缩成极小的缩略图(比如64x64像素),再用高压缩率格式保存后转Base64。生成的字符串大小只有原图Base64的1/10甚至更小,但仍能承载足够LLM做判断的视觉信息。
  • 代码示例:
    from PIL import Image
    import base64
    from io import BytesIO
    
    def get_compressed_thumbnail_base64(img_path, max_dim=64):
        with Image.open(img_path) as img:
            # 按比例缩放到指定最大尺寸
            img.thumbnail((max_dim, max_dim))
            # 内存中保存为高压缩率的JPEG/PNG
            buffer = BytesIO()
            # 照片用JPEG(质量30-50足够),透明图用PNG
            img_format = "JPEG" if img.mode in ["RGB", "L"] else "PNG"
            img.save(buffer, format=img_format, quality=40)
            # 转Base64字符串
            return base64.b64encode(buffer.getvalue()).decode("utf-8")
    
    # 使用示例
    llm_visual_input = get_compressed_thumbnail_base64("./my_target_image.png")
    
  • 小技巧:根据图像类型选格式——照片用JPEG压缩率更高,带透明通道的图用PNG,能进一步缩小Base64的长度。

3. 本地文件服务代理(适用于本地部署的LLM)

如果你在本地运行LLM(而非调用云端API),可以搭一个极简的静态文件服务,让LLM通过URL直接访问图像文件。

  • 核心思路:把图像放在指定目录,用Python内置的http.server启动轻量服务,然后把本地URL传给LLM,支持URL图像读取的LLM就能直接获取图像数据。
  • 快速实现:
    1. 创建temp_image_storage文件夹,把图像移到该目录下
    2. 终端启动服务:
      python -m http.server 8000 --directory ./temp_image_storage
      
    3. 传给LLM的内容:http://localhost:8000/my_target_image.png
  • 注意:这个方案只适用于本地LLM,云端LLM无法访问你的本地网络服务,需谨慎使用。

总结建议:如果LLM只需要基于图像元数据做决策,优先选「特征摘要+ID映射」方案;如果需要LLM做基础视觉判断,用「压缩缩略图Base64」;本地部署LLM的场景,「文件服务代理」是最省心的选择。

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 08:38:08