You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python删除COCO标注JSON中category_id为7的条目失效问题

核心问题原因

原代码存在4个直接导致功能失效的错误:

  • 索引起始值错误:循环使用range(1, len(data["annotations"]))生成索引,直接跳过了列表索引为0的第一个元素,示例中第一条category_id=7的标注恰好位于索引0,从遍历起始就被漏检。
  • 动态删列表的逻辑错误:for i in range(...)的索引序列在循环启动时就已经固定生成,循环内部手动执行i = i - 1不会改变下一轮迭代的i取值。删除元素后列表后续元素会整体前移,很容易出现连续待删条目漏检、索引越界的问题。
  • 缺少结果持久化逻辑:代码仅在内存中修改加载的JSON数据,没有将修改后的内容写回磁盘文件,即使内存中数据修改正确,原标注文件也不会发生任何变化。
  • 未同步清理分类表:仅删除annotations下的条目时,categories字段中残留的无效分类项会导致后续标注解析出错。
正确实现方案

不要在遍历原列表时直接执行删除操作,采用新建列表追加符合条件条目的方式从根源避免索引错位问题,同时支持批量修改category ID、同步清理无效分类、结果写回文件的完整需求:

import json

annotations_test = 'test_annotations/_annotations.coco_test.json'
annotations_train = 'train/_annotations.coco.json'

def process_coco_annotation(ann_path, keep_category_ids, category_id_map=None, save_path=None):
    """
    处理COCO格式标注文件
    :param ann_path: 原始标注文件路径
    :param keep_category_ids: 列表,传入需要保留的category_id,不在列表内的标注会被删除
    :param category_id_map: 字典,可选,格式为{旧category_id: 新category_id},用于批量修改标签ID
    :param save_path: 处理后文件的保存路径,默认覆盖原文件
    """
    if save_path is None:
        save_path = ann_path
    if category_id_map is None:
        category_id_map = {}

    # 加载原始标注
    with open(ann_path, "r", encoding="utf-8") as f:
        data = json.load(f)

    # 过滤标注条目,同步修改需要调整的category_id
    filtered_annotations = []
    for ann in data["annotations"]:
        cat_id = ann["category_id"]
        if cat_id not in keep_category_ids:
            continue
        if cat_id in category_id_map:
            ann["category_id"] = category_id_map[cat_id]
        filtered_annotations.append(ann)
    data["annotations"] = filtered_annotations

    # 同步清理分类表,更新分类ID
    filtered_categories = []
    for cat in data["categories"]:
        old_id = cat["id"]
        if old_id not in keep_category_ids:
            continue
        if old_id in category_id_map:
            cat["id"] = category_id_map[old_id]
        filtered_categories.append(cat)
    data["categories"] = filtered_categories

    # 将修改后的结果写回文件
    with open(save_path, "w", encoding="utf-8") as f:
        json.dump(data, f, indent=2, ensure_ascii=False)
    
    return data


# 调用示例:删除category_id=7的所有条目,保留id为0、3的类别,同时将原id为3的类别修改为id=1
process_coco_annotation(
    ann_path=annotations_test,
    keep_category_ids=[0, 3],
    category_id_map={3: 1}
)

操作提示:如果不需要修改category ID,调用函数时不传category_id_map参数即可;如果需要处理训练集标注,把ann_path参数换成annotations_train对应的路径就行。

原代码漏删的典型场景

当两个待删除的标注连续排列时,删除索引为i的元素后,原本在i+1位置的待删元素会前移到i位置,但下一轮循环i会自动自增为i+1,直接跳过这个前移的待删元素;再加上循环从索引1开始,所有位于索引0的待删条目会100%漏删。

内容的提问来源于stack exchange,提问作者philip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.05 16:15:46