求助:如何将自定义图片标注JSON与图片导入FiftyOne
解决FiftyOne导入自定义标注JSON的问题
问题描述
现有图片数据集和对应的标注JSON文件,尝试用FiftyOne的from_json或from_dir方法导入时报错,自行编写的代码运行耗时过长无法完成,代码如下:
import glob import fiftyone as fo import json images_patt = "./data/*.jpg" # Ex: load your custom label format with open('./Waste_Plastics_result.json', 'r') as f: annotations = json.load(f) # Create samples for your data samples = [] for filepath in glob.glob(images_patt): sample = fo.Sample(filepath=filepath) # Convert detections to FiftyOne format detections = [] for obj in annotations: width = obj["width"] height = obj["height"] #id = obj["id"] file_name = obj["file_name"] #image_id = obj["image_id"] category_id = obj["category_id"] metainfo_id = obj["metainfo_id"] # Bounding box coordinates should be relative values # in [0, 1] in the following format: # [top-left-x, top-left-y, width, height] bbox = obj["bbox"] #ignore = obj["ignore"] #iscrowd = obj["iscrowd"] area = obj["area"] detections.append( fo.Detection(width=width, height=height, file_name=file_name, category_id=category_id, metainfo_id=metainfo_id, bbox=bbox, area=area) ) # Store detections in a field name of your choice sample["ground_truth"] = fo.Detections(detections=detections) samples.append(sample) # Create dataset dataset = fo.Dataset.from_images_dir("./data") dataset.add_samples(samples) # FiftyOne session session = fo.launch_app(dataset) session.wait()
代码问题分析
- 嵌套循环逻辑错误:遍历每张图片时,把所有标注都添加到当前样本中,导致每个样本绑定了全量标注,数据冗余且计算量暴增,这是耗时过长的核心原因。
- 重复创建数据集:先通过
Dataset.from_images_dir导入所有图片,又调用add_samples重复添加样本,造成数据重复加载。 - BBox格式不匹配:标注中的bbox是像素级绝对坐标,但FiftyOne的
fo.Detection要求bbox为相对坐标(范围[0,1]),需要基于图片宽高转换。 - 无效参数传入:
fo.Detection不需要width、height、file_name这些图片属性,传入后会造成冗余数据。
修正后的代码
import fiftyone as fo import json import os # 1. 加载标注并建立图片名到标注的映射 annotations_path = "./Waste_Plastics_result.json" images_dir = "./data" with open(annotations_path, 'r') as f: annotations = json.load(f) # 按图片文件名分组标注 img_annot_map = {} for ann in annotations: img_name = ann["file_name"] if img_name not in img_annot_map: img_annot_map[img_name] = [] img_annot_map[img_name].append(ann) # 2. 遍历图片,创建样本 samples = [] for img_name in os.listdir(images_dir): if not img_name.endswith(".jpg"): continue img_path = os.path.join(images_dir, img_name) sample = fo.Sample(filepath=img_path) # 获取当前图片的标注 if img_name not in img_annot_map: # 无标注的图片也加入数据集,ground_truth为空 sample["ground_truth"] = fo.Detections(detections=[]) samples.append(sample) continue # 转换标注为FiftyOne格式 detections = [] for ann in img_annot_map[img_name]: # 获取图片宽高(从标注中读取,或从图片文件读取) img_width = ann["width"] img_height = ann["height"] # 转换绝对bbox为相对坐标:[x, y, w, h] -> [x/img_width, y/img_height, w/img_width, h/img_height] bbox_abs = ann["bbox"] bbox_rel = [ bbox_abs[0] / img_width, bbox_abs[1] / img_height, bbox_abs[2] / img_width, bbox_abs[3] / img_height ] # 创建Detection对象,只保留必要字段 det = fo.Detection( category_id=ann["category_id"], bbox=bbox_rel, area=ann["area"] # 如果有类别名称,可以添加label字段,比如label=category_names[ann["category_id"]] ) detections.append(det) sample["ground_truth"] = fo.Detections(detections=detections) samples.append(sample) # 3. 创建数据集并启动App dataset = fo.Dataset("waste_plastics_dataset") dataset.add_samples(samples) session = fo.launch_app(dataset) session.wait()
额外说明
- 如果标注中包含类别名称映射(比如
categories字段),可以添加label字段到fo.Detection中,提升可视化效果。 - 若图片宽高在标注中不存在,可以用PIL库读取图片获取:
from PIL import Image with Image.open(img_path) as img: img_width, img_height = img.size
内容的提问来源于stack exchange,提问作者user21840384
相关产品推荐
相关产品推荐

