深度学习中图像分辨率过大致内存溢出,如何分割图像并对应导出标注?
解决图像分割与标注匹配的实现方案
核心思路
将大图像按固定尺寸(带重叠)裁剪为小图,同时解析JSON标注,把每个目标映射到对应的小图中,转换为TXT格式的本地标注,确保小图与标注一一对应。
具体实现步骤(Python示例)
1. 准备依赖与参数配置
先导入所需库,设置裁剪尺寸、重叠率、保存路径等参数:
import cv2 import json import os from math import ceil # 配置参数 CROP_SIZE = (512, 512) # 小图尺寸(宽, 高) OVERLAP_RATIO = 0.1 # 重叠比例,避免边缘目标被截断 IMAGE_PATH = "path/to/your/image.jpg" JSON_PATH = "path/to/your/annotations.json" OUTPUT_IMG_DIR = "cropped_images" OUTPUT_TXT_DIR = "cropped_labels" # 创建输出目录 os.makedirs(OUTPUT_IMG_DIR, exist_ok=True) os.makedirs(OUTPUT_TXT_DIR, exist_ok=True)
2. 加载图像与解析JSON标注
假设你的JSON标注格式为COCO风格(每个目标包含bbox和category_id),解析标注并转换为统一的(x_min, y_min, x_max, y_max)格式:
# 加载原图 img = cv2.imread(IMAGE_PATH) img_h, img_w = img.shape[:2] # 解析JSON标注 with open(JSON_PATH, "r") as f: annotations = json.load(f)["annotations"] # 假设COCO格式,按需调整 # 转换标注为(x_min, y_min, x_max, y_max),并存储 label_list = [] for ann in annotations: x, y, w, h = ann["bbox"] x_min = x y_min = y x_max = x + w y_max = y + h category_id = ann["category_id"] label_list.append((x_min, y_min, x_max, y_max, category_id))
3. 裁剪图像并匹配标注
遍历图像的每个裁剪区域,计算该区域内的有效标注,转换为小图的相对坐标:
# 计算裁剪步长 step_w = int(CROP_SIZE[0] * (1 - OVERLAP_RATIO)) step_h = int(CROP_SIZE[1] * (1 - OVERLAP_RATIO)) # 遍历所有裁剪区域 crop_idx = 0 for y_start in range(0, img_h, step_h): for x_start in range(0, img_w, step_w): # 计算裁剪区域的实际结束坐标(避免超出图像边界) x_end = min(x_start + CROP_SIZE[0], img_w) y_end = min(y_start + CROP_SIZE[1], img_h) # 裁剪小图 cropped_img = img[y_start:y_end, x_start:x_end] crop_h, crop_w = cropped_img.shape[:2] # 匹配当前裁剪区域的标注 current_labels = [] for (x_min, y_min, x_max, y_max, cls_id) in label_list: # 计算目标与裁剪区域的交集 inter_x_min = max(x_min, x_start) inter_y_min = max(y_min, y_start) inter_x_max = min(x_max, x_end) inter_y_max = min(y_max, y_end) # 跳过无交集或交集过小的目标(可调整阈值) if inter_x_max <= inter_x_min or inter_y_max <= inter_y_min: continue inter_area = (inter_x_max - inter_x_min) * (inter_y_max - inter_y_min) obj_area = (x_max - x_min) * (y_max - y_min) if inter_area / obj_area < 0.5: # 保留超过50%在小图内的目标 continue # 转换为小图内的相对坐标(YOLO格式:x_center, y_center, w, h,归一化到0-1) rel_x_center = (inter_x_min + inter_x_max) / 2 - x_start rel_y_center = (inter_y_min + inter_y_max) / 2 - y_start rel_w = inter_x_max - inter_x_min rel_h = inter_y_max - inter_y_min # 归一化 rel_x_center /= crop_w rel_y_center /= crop_h rel_w /= crop_w rel_h /= crop_h current_labels.append(f"{cls_id} {rel_x_center:.6f} {rel_y_center:.6f} {rel_w:.6f} {rel_h:.6f}") # 保存小图和标注TXT crop_name = f"crop_{crop_idx:04d}_x{x_start}_y{y_start}" cv2.imwrite(os.path.join(OUTPUT_IMG_DIR, f"{crop_name}.jpg"), cropped_img) if current_labels: with open(os.path.join(OUTPUT_TXT_DIR, f"{crop_name}.txt"), "w") as f: f.write("\n".join(current_labels)) crop_idx += 1
4. 自定义调整要点
- 标注格式适配:如果你的JSON不是COCO格式,只需修改
label_list的生成逻辑,提取目标的边界框和类别即可。 - 重叠率与阈值:如果边缘目标较多,可提高
OVERLAP_RATIO;调整inter_area / obj_area的阈值,决定保留多少比例的目标。 - 小图尺寸:根据你的模型输入尺寸设置
CROP_SIZE,确保与模型要求匹配。
内容的提问来源于stack exchange,提问作者Mumu
相关产品推荐
相关产品推荐

