You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

深度学习中图像分辨率过大致内存溢出,如何分割图像并对应导出标注?

解决图像分割与标注匹配的实现方案

核心思路

将大图像按固定尺寸(带重叠)裁剪为小图,同时解析JSON标注,把每个目标映射到对应的小图中,转换为TXT格式的本地标注,确保小图与标注一一对应。

具体实现步骤(Python示例)

1. 准备依赖与参数配置

先导入所需库,设置裁剪尺寸、重叠率、保存路径等参数:

import cv2
import json
import os
from math import ceil

# 配置参数
CROP_SIZE = (512, 512)  # 小图尺寸(宽, 高)
OVERLAP_RATIO = 0.1     # 重叠比例,避免边缘目标被截断
IMAGE_PATH = "path/to/your/image.jpg"
JSON_PATH = "path/to/your/annotations.json"
OUTPUT_IMG_DIR = "cropped_images"
OUTPUT_TXT_DIR = "cropped_labels"

# 创建输出目录
os.makedirs(OUTPUT_IMG_DIR, exist_ok=True)
os.makedirs(OUTPUT_TXT_DIR, exist_ok=True)

2. 加载图像与解析JSON标注

假设你的JSON标注格式为COCO风格(每个目标包含bbox和category_id),解析标注并转换为统一的(x_min, y_min, x_max, y_max)格式:

# 加载原图
img = cv2.imread(IMAGE_PATH)
img_h, img_w = img.shape[:2]

# 解析JSON标注
with open(JSON_PATH, "r") as f:
    annotations = json.load(f)["annotations"]  # 假设COCO格式,按需调整

# 转换标注为(x_min, y_min, x_max, y_max),并存储
label_list = []
for ann in annotations:
    x, y, w, h = ann["bbox"]
    x_min = x
    y_min = y
    x_max = x + w
    y_max = y + h
    category_id = ann["category_id"]
    label_list.append((x_min, y_min, x_max, y_max, category_id))

3. 裁剪图像并匹配标注

遍历图像的每个裁剪区域,计算该区域内的有效标注,转换为小图的相对坐标:

# 计算裁剪步长
step_w = int(CROP_SIZE[0] * (1 - OVERLAP_RATIO))
step_h = int(CROP_SIZE[1] * (1 - OVERLAP_RATIO))

# 遍历所有裁剪区域
crop_idx = 0
for y_start in range(0, img_h, step_h):
    for x_start in range(0, img_w, step_w):
        # 计算裁剪区域的实际结束坐标(避免超出图像边界)
        x_end = min(x_start + CROP_SIZE[0], img_w)
        y_end = min(y_start + CROP_SIZE[1], img_h)
        
        # 裁剪小图
        cropped_img = img[y_start:y_end, x_start:x_end]
        crop_h, crop_w = cropped_img.shape[:2]
        
        # 匹配当前裁剪区域的标注
        current_labels = []
        for (x_min, y_min, x_max, y_max, cls_id) in label_list:
            # 计算目标与裁剪区域的交集
            inter_x_min = max(x_min, x_start)
            inter_y_min = max(y_min, y_start)
            inter_x_max = min(x_max, x_end)
            inter_y_max = min(y_max, y_end)
            
            # 跳过无交集或交集过小的目标(可调整阈值)
            if inter_x_max <= inter_x_min or inter_y_max <= inter_y_min:
                continue
            inter_area = (inter_x_max - inter_x_min) * (inter_y_max - inter_y_min)
            obj_area = (x_max - x_min) * (y_max - y_min)
            if inter_area / obj_area < 0.5:  # 保留超过50%在小图内的目标
                continue
            
            # 转换为小图内的相对坐标(YOLO格式:x_center, y_center, w, h,归一化到0-1)
            rel_x_center = (inter_x_min + inter_x_max) / 2 - x_start
            rel_y_center = (inter_y_min + inter_y_max) / 2 - y_start
            rel_w = inter_x_max - inter_x_min
            rel_h = inter_y_max - inter_y_min
            
            # 归一化
            rel_x_center /= crop_w
            rel_y_center /= crop_h
            rel_w /= crop_w
            rel_h /= crop_h
            
            current_labels.append(f"{cls_id} {rel_x_center:.6f} {rel_y_center:.6f} {rel_w:.6f} {rel_h:.6f}")
        
        # 保存小图和标注TXT
        crop_name = f"crop_{crop_idx:04d}_x{x_start}_y{y_start}"
        cv2.imwrite(os.path.join(OUTPUT_IMG_DIR, f"{crop_name}.jpg"), cropped_img)
        if current_labels:
            with open(os.path.join(OUTPUT_TXT_DIR, f"{crop_name}.txt"), "w") as f:
                f.write("\n".join(current_labels))
        
        crop_idx += 1

4. 自定义调整要点

  • 标注格式适配:如果你的JSON不是COCO格式,只需修改label_list的生成逻辑,提取目标的边界框和类别即可。
  • 重叠率与阈值:如果边缘目标较多,可提高OVERLAP_RATIO;调整inter_area / obj_area的阈值,决定保留多少比例的目标。
  • 小图尺寸:根据你的模型输入尺寸设置CROP_SIZE,确保与模型要求匹配。

内容的提问来源于stack exchange,提问作者Mumu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 04:33:10