You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Segment Anything Model裁剪图片时遇类型错误求助

解决SAM分割后裁剪图片的TypeError问题

错误原因

mask_generator.generate()返回的是字典列表,每个字典包含segmentation(掩码数组)、bbox、area等字段。你直接用result[0](字典对象)和0.5做比较,自然会抛出TypeError。

解决方案

需要从每个掩码字典中提取segmentation字段(布尔型数组),转换为OpenCV可处理的格式后再提取轮廓,同时可以通过面积过滤排除小噪声区域,最终完成裁剪。

修正后的完整代码

IMAGE_PATH = "Akher-Saa-001-1934_Page_05.jpg"
import cv2
from segment_anything import SamAutomaticMaskGenerator
import torch
from segment_anything import sam_model_registry
import supervision as sv


DEVICE = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')
MODEL_TYPE = "vit_h"

sam = sam_model_registry[MODEL_TYPE](checkpoint="/home/f4/Afifi/sam_vit_h_4b8939.pth")
sam.to(device=DEVICE)
mask_generator = SamAutomaticMaskGenerator(sam)

image_bgr = cv2.imread(IMAGE_PATH)
image_rgb = cv2.cvtColor(image_bgr, cv2.COLOR_BGR2RGB)
result = mask_generator.generate(image_rgb)

mask_annotator = sv.MaskAnnotator(color_lookup = sv.ColorLookup.INDEX)
detections = sv.Detections.from_sam(result)
annotated_image = mask_annotator.annotate(image_bgr, detections)

annotated_image_path = "annotated_image.jpg"
cv2.imwrite(annotated_image_path, annotated_image)
print(f"Annotated image saved at: {annotated_image_path}")

# 遍历所有分割出的掩码
for mask_idx, mask_dict in enumerate(result):
    # 提取掩码数组,转换为OpenCV可用的8位单通道格式
    mask = mask_dict['segmentation'].astype('uint8') * 255
    
    # 过滤小面积区域,避免噪声(可根据需求调整阈值)
    if mask_dict['area'] < 1000:
        continue
    
    # 提取轮廓
    contours, _ = cv2.findContours(mask, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
    
    for contour_idx, contour in enumerate(contours):
        x, y, w, h = cv2.boundingRect(contour)
        # 裁剪区域
        cropped_region = image_bgr[y:y + h, x:x + w]
        # 格式化输出路径(修正原代码的字符串格式化错误)
        crop_output_path = f"cropped_region_{mask_idx}_{contour_idx}.jpg"
        cv2.imwrite(crop_output_path, cropped_region)
        print(f"Cropped region saved at: {crop_output_path}")

关键修改点

  1. 提取掩码数组:用mask_dict['segmentation']获取真实的掩码数据,替代直接使用result[0]
  2. 格式转换:将布尔数组转为uint8类型并乘以255,符合OpenCV对轮廓提取的图像格式要求
  3. 噪声过滤:通过mask_dict['area']设置阈值,跳过过小的分割区域,避免裁剪无效内容
  4. 路径修复:用f-string格式化输出路径,确保文件名正确包含索引

内容的提问来源于stack exchange,提问作者Mohamed Mostafa Afify

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 14:02:50