You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为4000张含多目标的图片进行自定义多边形标注?

解决方案:自定义物体多边形自动标注并输出JSON

一、基于预训练模型微调+标注工具的流程

1. 准备少量手动标注样本

先手动标注50-100张包含自定义目标的图片,用多边形标注,存成JSON格式(可以用LabelMe这类工具,它原生支持多边形和JSON输出)。这部分是给模型做微调的训练数据。

2. 微调实例分割模型

选一个开源的实例分割模型(比如Mask R-CNN、YOLOv8-seg),用你准备的小样本数据集微调:

  • 把LabelMe导出的JSON转换成模型需要的格式(比如COCO格式,很多模型都支持),可以写个简单的Python脚本批量转换:
    # 示例:LabelMe JSON转COCO格式核心逻辑
    import json
    import os
    
    def labelme2coco(labelme_json_dir, output_json):
        coco_data = {"images": [], "annotations": [], "categories": []}
        # 添加自定义类别,比如你的类别叫"custom_obj"
        coco_data["categories"].append({"id": 1, "name": "custom_obj", "supercategory": "none"})
        
        img_id = 1
        ann_id = 1
        for json_file in os.listdir(labelme_json_dir):
            if not json_file.endswith(".json"):
                continue
            with open(os.path.join(labelme_json_dir, json_file), "r") as f:
                data = json.load(f)
            # 处理图片信息
            img_info = {
                "id": img_id,
                "file_name": data["imagePath"],
                "height": data["imageHeight"],
                "width": data["imageWidth"]
            }
            coco_data["images"].append(img_info)
            # 处理标注信息
            for shape in data["shapes"]:
                if shape["label"] != "custom_obj":
                    continue
                # 提取多边形点
                points = shape["points"]
                # 转换为COCO格式的segmentation和bbox
                segmentation = [sum(points, [])]
                x_coords = [p[0] for p in points]
                y_coords = [p[1] for p in points]
                bbox = [min(x_coords), min(y_coords), max(x_coords)-min(x_coords), max(y_coords)-min(y_coords)]
                ann_info = {
                    "id": ann_id,
                    "image_id": img_id,
                    "category_id": 1,
                    "segmentation": segmentation,
                    "bbox": bbox,
                    "area": bbox[2]*bbox[3],
                    "iscrowd": 0
                }
                coco_data["annotations"].append(ann_info)
                ann_id +=1
            img_id +=1
        with open(output_json, "w") as f:
            json.dump(coco_data, f)
    
  • 用微调后的模型批量处理剩下的图片,生成每个目标的多边形掩码/点坐标。

3. 模型结果转JSON并校验

  • 写脚本把模型输出的分割结果转换成你需要的JSON格式,比如和LabelMe一致的结构,方便后续校验:
    # 示例:模型分割结果转LabelMe风格JSON
    import json
    import cv2
    import os
    
    def seg_result2labelme(img_path, seg_points, label, output_json):
        img = cv2.imread(img_path)
        height, width = img.shape[:2]
        data = {
            "version": "5.0.1",
            "flags": {},
            "shapes": [
                {
                    "label": label,
                    "points": seg_points,
                    "group_id": None,
                    "shape_type": "polygon",
                    "flags": {}
                }
            ],
            "imagePath": os.path.basename(img_path),
            "imageData": None,
            "imageHeight": height,
            "imageWidth": width
        }
        with open(output_json, "w") as f:
            json.dump(data, f)
    
  • 用LabelMe打开自动生成的JSON文件,快速检查和修正错误标注,比全手动快很多。

二、无代码/轻代码替代方案

如果不想自己微调模型,可以用支持自定义实例分割的标注工具:

  • LabelStudio:支持导入自定义数据集,集成ML后端,上传少量标注样本让模型学习后,就能批量自动标注剩下的图,可直接导出JSON格式的多边形标注。
  • CVAT:支持自定义类别和多边形标注,自带自动标注功能,能导出COCO或自定义JSON格式,适合大规模数据集处理。

三、关键注意事项

  • 微调模型时,小样本要覆盖自定义目标的不同场景(比如不同角度、光照、背景),提升自动标注准确率。
  • 自动标注后一定要抽样校验,修正错误,避免数据集出现偏差。
  • 如果自定义目标特征独特,可选用Few-Shot实例分割模型(比如FewShotDet),仅需少量标注样本就能得到不错的效果。

内容的提问来源于stack exchange,提问作者nobody

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 03:35:10