Mask RCNN如何适配单图单JSON标注?修改函数还是合并文件?
适配单图单JSON标注到Mask RCNN的解决方案
针对你用Mask RCNN做自定义目标检测时,遇到的单图单JSON标注和官方要求不匹配的问题,我从你的两个核心疑问入手,给你详细的落地方案:
疑问2:是否需要合并所有JSON?合并后的格式符合要求吗?
我非常推荐你把所有单图对应的JSON标注合并成一个总文件,这不仅贴合Mask RCNN官方的设计逻辑,还能避免频繁读取大量小文件的性能损耗,后续数据加载也会更顺畅。
合并后的标注格式需要参考官方数据集(比如Balloon数据集)的结构,大致要包含images、annotations、categories三个核心部分,示例结构如下:
{ "images": [ { "id": 1, "file_name": "你的图像文件名.jpg", "width": 图像宽度(像素), "height": 图像高度(像素) }, // 其他所有图像的信息都放在这里 ], "annotations": [ { "id": 1, "image_id": 1, // 和上面images里的id对应 "category_id": 1, // 和categories里的id对应 "segmentation": [[35.70,18.38,60.57,15.68,...]], // 多边形点展开成一维数组;圆形要转成近似多边形 "bbox": [x_min, y_min, 框宽, 框高], // 目标的外接矩形 "area": 目标区域的面积值, "iscrowd": 0 // 固定0就行,你的场景不需要crowd标注 }, // 所有图像的标注都放在这里 ], "categories": [ { "id": 1, "name": "anchor" // 你的目标类别 } ] }
针对你标注里的两种形状,处理方式要注意:
- 多边形(polygon):直接把标注里的
points列表按顺序展开成一维数组,放到segmentation字段即可。 - 圆形(circle):Mask RCNN不支持原生的圆形mask,你得把圆形转换成近似的多边形(比如用32个点拟合圆周就足够精确了),再把这些点的坐标放到
segmentation里;同时计算圆形的外接矩形作为bbox,面积用圆的面积公式πr²计算就行。
你可以写个简单的Python脚本批量处理:遍历所有图像和对应的JSON,收集图像的宽高信息,转换标注形状,最后把所有内容整合成上面的格式。
疑问1:如何修改export_boxes和load_mask函数?
我分两种场景给你示例代码,优先推荐用合并后的方案:
场景1:已经合并为单个总JSON(推荐)
这种情况下,参考官方Balloon数据集的实现,做少量调整就能适配:
1. 修改load_mask函数
这个函数的作用是返回当前图像对应的mask数组和类别ID,核心逻辑是根据图像ID匹配标注,再生成mask:
def load_mask(self, image_id): # 获取当前图像的基本信息 info = self.image_info[image_id] # 从总标注里筛选出当前图像的所有标注 annotations = [a for a in self.dataset['annotations'] if a['image_id'] == info['id']] count = len(annotations) # 初始化mask数组(形状:图像高×图像宽×标注数量) mask = np.zeros([info['height'], info['width'], count], dtype=np.uint8) class_ids = np.zeros(count, dtype=np.int32) for i, ann in enumerate(annotations): # 把segmentation的一维数组转成二维点集 polygon = np.array(ann['segmentation']).reshape(-1, 2) # 用skimage生成多边形mask rr, cc = skimage.draw.polygon(polygon[:, 1], polygon[:, 0]) mask[rr, cc, i] = 1 # 设置类别ID(这里假设你的类别anchor对应的id是1) class_ids[i] = self.class_names.index(ann['category_id']) return mask.astype(np.bool), class_ids.astype(np.int32)
2. 修改export_boxes函数
这个函数用来导出目标的边界框信息,直接从标注里提取bbox即可:
def export_boxes(self, dataset): boxes = [] for image_id in dataset.image_ids: info = dataset.image_info[image_id] annotations = [a for a in dataset.dataset['annotations'] if a['image_id'] == info['id']] for ann in annotations: bbox = ann['bbox'] # 把[x_min, y_min, 宽, 高]转成[x_min, y_min, x_max, y_max]格式 boxes.append({ 'image_id': image_id, 'file_name': info['file_name'], 'bbox': [bbox[0], bbox[1], bbox[0]+bbox[2], bbox[1]+bbox[3]], 'class_id': ann['category_id'], 'class_name': dataset.class_names[ann['category_id']] }) return boxes
场景2:不合并,保留单图单JSON(不推荐)
如果实在不想合并,那就要在函数里根据图像路径找到对应的JSON文件,再解析标注:
1. 修改load_mask函数
def load_mask(self, image_id): info = self.image_info[image_id] # 假设JSON和图像同目录,后缀替换成.json(比如xxx.jpg对应xxx.json) json_path = os.path.splitext(info['path'])[0] + '.json' with open(json_path, 'r') as f: annotations = json.load(f) count = len(annotations) mask = np.zeros([info['height'], info['width'], count], dtype=np.uint8) class_ids = np.zeros(count, dtype=np.int32) for i, ann in enumerate(annotations): # 设置类别ID class_ids[i] = self.class_names.index(ann['label']) if ann['shape_type'] == 'polygon': points = np.array(ann['points']).reshape(-1, 2) rr, cc = skimage.draw.polygon(points[:, 1], points[:, 0]) mask[rr, cc, i] = 1 elif ann['shape_type'] == 'circle': # 把圆形转成32个点的多边形 center_x, center_y = ann['points'][0] # 计算半径:两点之间的距离 radius = np.linalg.norm(np.array(ann['points'][1]) - np.array(ann['points'][0])) # 生成圆周上的点 theta = np.linspace(0, 2*np.pi, 32) x = center_x + radius * np.cos(theta) y = center_y + radius * np.sin(theta) polygon = np.stack([x, y], axis=1) rr, cc = skimage.draw.polygon(polygon[:, 1], polygon[:, 0]) mask[rr, cc, i] = 1 return mask.astype(np.bool), class_ids.astype(np.int32)
2. 修改export_boxes函数
def export_boxes(self, dataset): boxes = [] for image_id in dataset.image_ids: info = dataset.image_info[image_id] json_path = os.path.splitext(info['path'])[0] + '.json' with open(json_path, 'r') as f: annotations = json.load(f) for ann in annotations: if ann['shape_type'] == 'polygon': points = np.array(ann['points']) # 计算多边形的外接矩形 x_min = np.min(points[:, 0]) y_min = np.min(points[:, 1]) x_max = np.max(points[:, 0]) y_max = np.max(points[:, 1]) elif ann['shape_type'] == 'circle': center_x, center_y = ann['points'][0] radius = np.linalg.norm(np.array(ann['points'][1]) - np.array(ann['points'][0])) x_min = center_x - radius y_min = center_y - radius x_max = center_x + radius y_max = center_y + radius boxes.append({ 'image_id': image_id, 'file_name': info['file_name'], 'bbox': [x_min, y_min, x_max, y_max], 'class_id': dataset.class_names.index(ann['label']), 'class_name': ann['label'] }) return boxes
最后给你个小提醒:不管用哪种方案,一定要确保标注的坐标是像素单位,和图像的实际分辨率匹配;圆形转多边形时,点数不用太多,32个点就足够平衡精度和效率了。
内容的提问来源于stack exchange,提问作者y_1234
相关产品推荐
相关产品推荐

