Python遍历JSON的annotations键时删除无效标注行的方法
清理JSON图像标注中的无效条目
我手头有个存图像标注数据的大型JSON文件,需要遍历其中的annotations列表,删掉两类无效标注:一类是segmentation字段值为[[]]的,另一类是bbox为空数组的。
最初尝试的无效代码
一开始我写了这段代码,但根本删不掉目标条目:
import json # 打开JSON文件 f = open('annotations.json') # 加载为字典格式 data = json.load(f) # 遍历标注列表 for i in data['annotations']: if i['segmentation'] == [[]]: print(i['segmentation']) del i # 关闭文件 f.close()
问题出在del i只是删除了循环变量的引用,并没有从data['annotations']这个列表里移除对应元素,原数据根本没被修改。
无效标注条目示例
这类无效条目长这样:
{"iscrowd":0,"image_id":32,"bbox":[],"segmentation":[[]],"category_id":2,"id":339,"area":0}
最终可行的代码
后来调整代码实现了需求:
import json # 打开JSON文件 f = open('annotations.json') # 加载为字典格式 data = json.load(f) # 关闭文件 f.close() # 遍历标注列表 count = 0 for key in data['annotations']: count +=1 if key['segmentation'] == [[]]: print(key['segmentation']) data["annotations"].pop(count) if key['bbox'] == []: data["annotations"].pop(count) with open("newannotations.json", "w") as json_file: json.dump(data, json_file)
更稳妥的优化写法
上面的代码用pop(count)可能会因为列表长度变化导致索引错位(比如连续两个无效条目时,第二个可能漏删),更简洁安全的方式是用列表推导式直接生成过滤后的列表:
import json # 用with语句自动管理文件,避免忘记关闭 with open('annotations.json', 'r') as f: data = json.load(f) # 过滤掉两类无效标注 data['annotations'] = [ ann for ann in data['annotations'] if ann['segmentation'] != [[]] and ann['bbox'] != [] ] # 写入新文件,加indent让JSON格式更易读 with open("newannotations.json", "w") as json_file: json.dump(data, json_file, indent=2)
内容的提问来源于stack exchange,提问作者devman3211
相关产品推荐
相关产品推荐

