You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历JSON的annotations键时删除无效标注行的方法

清理JSON图像标注中的无效条目

我手头有个存图像标注数据的大型JSON文件,需要遍历其中的annotations列表,删掉两类无效标注:一类是segmentation字段值为[[]]的,另一类是bbox为空数组的。

最初尝试的无效代码

一开始我写了这段代码,但根本删不掉目标条目:

import json
  
# 打开JSON文件
f = open('annotations.json')
  
# 加载为字典格式
data = json.load(f)
  
# 遍历标注列表
for i in data['annotations']:
    if i['segmentation'] == [[]]:
        print(i['segmentation'])
        del i
  
# 关闭文件
f.close()

问题出在del i只是删除了循环变量的引用,并没有从data['annotations']这个列表里移除对应元素,原数据根本没被修改。

无效标注条目示例

这类无效条目长这样:

{"iscrowd":0,"image_id":32,"bbox":[],"segmentation":[[]],"category_id":2,"id":339,"area":0}

最终可行的代码

后来调整代码实现了需求:

import json
  
# 打开JSON文件
f = open('annotations.json')
  
# 加载为字典格式
data = json.load(f)

# 关闭文件
f.close()

# 遍历标注列表
count = 0
for key in data['annotations']:
    count +=1
    if key['segmentation'] == [[]]:
        print(key['segmentation'])
        data["annotations"].pop(count)
    if key['bbox'] == []:
        data["annotations"].pop(count)

with open("newannotations.json", "w") as json_file:  
        json.dump(data, json_file)

更稳妥的优化写法

上面的代码用pop(count)可能会因为列表长度变化导致索引错位(比如连续两个无效条目时,第二个可能漏删),更简洁安全的方式是用列表推导式直接生成过滤后的列表:

import json

# 用with语句自动管理文件,避免忘记关闭
with open('annotations.json', 'r') as f:
    data = json.load(f)

# 过滤掉两类无效标注
data['annotations'] = [
    ann for ann in data['annotations']
    if ann['segmentation'] != [[]] and ann['bbox'] != []
]

# 写入新文件,加indent让JSON格式更易读
with open("newannotations.json", "w") as json_file:
    json.dump(data, json_file, indent=2)

内容的提问来源于stack exchange,提问作者devman3211

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.25 06:36:32