You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Meta Segment Anything模型时图像缩放后掩码对齐失败问题求助

解决4K图像SAM分割内存不足+掩码对齐问题

核心问题分析

你用cv2.resize将4K图像缩放到1920x1080后生成的掩码,尺寸与原图不匹配,直接套用到原图上必然出现对齐错误。要解决这个问题,需要在分割完成后,将所有掩码和对应的边界框按比例缩放回原图尺寸。

修改后的完整代码

import os
import torch
import cv2
import supervision as sv

HOME = os.getcwd()
print("HOME:", HOME)

CHECKPOINT_PATH = os.path.join(HOME, "weights", "sam_vit_h_4b8939.pth")
print(CHECKPOINT_PATH, "; exist:", os.path.isfile(CHECKPOINT_PATH))

DEVICE = torch.device('cuda:0' if torch.cuda.is_available() else 'cpu')
MODEL_TYPE = "vit_h"

from segment_anything import sam_model_registry, SamAutomaticMaskGenerator, SamPredictor

sam = sam_model_registry[MODEL_TYPE](checkpoint=CHECKPOINT_PATH).to(device=DEVICE)
mask_generator = SamAutomaticMaskGenerator(sam)

IMAGE_NAME = "prueba2.jpg"
IMAGE_PATH = os.path.join(HOME, "data", IMAGE_NAME)

# 读取原图并保存原始尺寸(注意cv2.shape是(高, 宽, 通道))
image_bgr = cv2.imread(IMAGE_PATH)
original_height, original_width = image_bgr.shape[:2]

# 缩放图像用于SAM处理
image_rgb = cv2.cvtColor(image_bgr, cv2.COLOR_BGR2RGB)
scaled_width, scaled_height = 1920, 1080
image_rgb_scaled = cv2.resize(image_rgb, (scaled_width, scaled_height))

# 生成掩码
sam_result = mask_generator.generate(image_rgb_scaled)

# 缩放掩码和边界框到原图尺寸
for mask_info in sam_result:
    # 缩放掩码,用最近邻插值保证二值属性
    mask_info['segmentation'] = cv2.resize(
        mask_info['segmentation'].astype('uint8'),
        (original_width, original_height),
        interpolation=cv2.INTER_NEAREST
    ).astype('bool')
    
    # 按比例更新边界框坐标
    width_ratio = original_width / scaled_width
    height_ratio = original_height / scaled_height
    x1, y1, w, h = mask_info['bbox']
    mask_info['bbox'] = [
        int(x1 * width_ratio),
        int(y1 * height_ratio),
        int(w * width_ratio),
        int(h * height_ratio)
    ]

# 标注并可视化
mask_annotator = sv.MaskAnnotator()
detections = sv.Detections.from_sam(sam_result=sam_result)
annotated_image = mask_annotator.annotate(scene=image_bgr.copy(), detections=detections)

sv.plot_images_grid(
    images=[image_bgr, annotated_image],
    grid_size=(1, 2),
    titles=['source image', 'segmented image']
)

关键修改说明

  1. 保存原图尺寸:读取原图后立即记录original_height和original_width,作为后续掩码缩放的目标尺寸
  2. 分离缩放变量:将缩放后的图像赋值给独立变量image_rgb_scaled,避免覆盖原图的RGB数据
  3. 掩码缩放处理:用INTER_NEAREST插值方式缩放二值掩码,确保掩码仅保留0/1的二值属性,避免模糊
  4. 边界框同步更新:根据宽高比例,将缩放后图像上的边界框坐标转换回原图坐标系,保证标注时的位置匹配

内容的提问来源于stack exchange,提问作者Daniel Sepulveda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 11:31:06