Python删除COCO标注JSON中category_id为7的条目失效问题
核心问题原因
原代码存在4个直接导致功能失效的错误:
- 索引起始值错误:循环使用
range(1, len(data["annotations"]))生成索引,直接跳过了列表索引为0的第一个元素,示例中第一条category_id=7的标注恰好位于索引0,从遍历起始就被漏检。 - 动态删列表的逻辑错误:
for i in range(...)的索引序列在循环启动时就已经固定生成,循环内部手动执行i = i - 1不会改变下一轮迭代的i取值。删除元素后列表后续元素会整体前移,很容易出现连续待删条目漏检、索引越界的问题。 - 缺少结果持久化逻辑:代码仅在内存中修改加载的JSON数据,没有将修改后的内容写回磁盘文件,即使内存中数据修改正确,原标注文件也不会发生任何变化。
- 未同步清理分类表:仅删除annotations下的条目时,
categories字段中残留的无效分类项会导致后续标注解析出错。
正确实现方案
不要在遍历原列表时直接执行删除操作,采用新建列表追加符合条件条目的方式从根源避免索引错位问题,同时支持批量修改category ID、同步清理无效分类、结果写回文件的完整需求:
import json annotations_test = 'test_annotations/_annotations.coco_test.json' annotations_train = 'train/_annotations.coco.json' def process_coco_annotation(ann_path, keep_category_ids, category_id_map=None, save_path=None): """ 处理COCO格式标注文件 :param ann_path: 原始标注文件路径 :param keep_category_ids: 列表,传入需要保留的category_id,不在列表内的标注会被删除 :param category_id_map: 字典,可选,格式为{旧category_id: 新category_id},用于批量修改标签ID :param save_path: 处理后文件的保存路径,默认覆盖原文件 """ if save_path is None: save_path = ann_path if category_id_map is None: category_id_map = {} # 加载原始标注 with open(ann_path, "r", encoding="utf-8") as f: data = json.load(f) # 过滤标注条目,同步修改需要调整的category_id filtered_annotations = [] for ann in data["annotations"]: cat_id = ann["category_id"] if cat_id not in keep_category_ids: continue if cat_id in category_id_map: ann["category_id"] = category_id_map[cat_id] filtered_annotations.append(ann) data["annotations"] = filtered_annotations # 同步清理分类表,更新分类ID filtered_categories = [] for cat in data["categories"]: old_id = cat["id"] if old_id not in keep_category_ids: continue if old_id in category_id_map: cat["id"] = category_id_map[old_id] filtered_categories.append(cat) data["categories"] = filtered_categories # 将修改后的结果写回文件 with open(save_path, "w", encoding="utf-8") as f: json.dump(data, f, indent=2, ensure_ascii=False) return data # 调用示例:删除category_id=7的所有条目,保留id为0、3的类别,同时将原id为3的类别修改为id=1 process_coco_annotation( ann_path=annotations_test, keep_category_ids=[0, 3], category_id_map={3: 1} )
操作提示:如果不需要修改category ID,调用函数时不传
category_id_map参数即可;如果需要处理训练集标注,把ann_path参数换成annotations_train对应的路径就行。
原代码漏删的典型场景
当两个待删除的标注连续排列时,删除索引为i的元素后,原本在i+1位置的待删元素会前移到i位置,但下一轮循环i会自动自增为i+1,直接跳过这个前移的待删元素;再加上循环从索引1开始,所有位于索引0的待删条目会100%漏删。
内容的提问来源于stack exchange,提问作者philip
相关产品推荐
相关产品推荐

