You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何下载Google Open Images V7用于YOLOv8训练?标签维度异常求助

解决Open Images V7下载与YOLOv8标签格式适配问题

问题根源

你使用的OID工具输出的标签格式不符合YOLOv8要求:YOLOv8需要归一化的中心坐标+宽高+类别索引,而该工具可能输出绝对坐标、类别名称,或边界框格式为(x1,y1,x2,y2)而非YOLO标准格式。

方案1:修正已下载的标签

如果已经完成图像和标签下载,可通过以下脚本批量转换格式:

步骤1:创建类别映射文件

在数据集根目录新建classes.txt,按顺序写入目标类别:

Apple
Orange

步骤2:运行标签转换脚本

import os
from PIL import Image

# 配置路径
dataset_dir = "./OID/Dataset/train"
classes_path = "./classes.txt"

# 加载类别到索引的映射
with open(classes_path, 'r') as f:
    classes = [line.strip() for line in f.readlines()]
class_map = {cls: idx for idx, cls in enumerate(classes)}

# 遍历每个类别文件夹
for cls_name in classes:
    cls_dir = os.path.join(dataset_dir, cls_name)
    old_label_dir = os.path.join(cls_dir, "Label")
    new_label_dir = os.path.join(cls_dir, "labels")
    os.makedirs(new_label_dir, exist_ok=True)
    
    # 处理每个标签文件
    for label_file in os.listdir(old_label_dir):
        if not label_file.endswith(".txt"):
            continue
        img_name = label_file.replace(".txt", ".jpg")
        img_path = os.path.join(cls_dir, img_name)
        
        # 获取图像尺寸用于归一化
        try:
            with Image.open(img_path) as img:
                img_w, img_h = img.size
        except:
            print(f"跳过损坏图像:{img_path}")
            continue
        
        # 转换标签格式
        new_lines = []
        with open(os.path.join(old_label_dir, label_file), 'r') as f:
            for line in f.readlines():
                parts = line.strip().split()
                if len(parts) < 5:
                    continue
                # OID标签格式:类别名 x_min y_min x_max y_max
                oid_cls = parts[0]
                x_min, y_min, x_max, y_max = map(float, parts[1:5])
                
                # 转换为YOLO格式:类别索引 x_center y_center width height(归一化到0-1)
                cls_idx = class_map.get(oid_cls, -1)
                if cls_idx == -1:
                    continue
                x_center = (x_min + x_max) / 2 / img_w
                y_center = (y_min + y_max) / 2 / img_h
                width = (x_max - x_min) / img_w
                height = (y_max - y_min) / img_h
                
                new_lines.append(f"{cls_idx} {x_center:.6f} {y_center:.6f} {width:.6f} {height:.6f}")
        
        # 保存转换后的标签
        with open(os.path.join(new_label_dir, label_file), 'w') as f:
            f.write('\n'.join(new_lines))

print("标签转换完成")

步骤3:整理数据集结构

调整为YOLOv8要求的标准结构:

dataset/
├── train/
│   ├── images/
│   │   ├── apple_001.jpg
│   │   ├── orange_001.jpg
│   │   └── ...
│   └── labels/
│       ├── apple_001.txt
│       ├── orange_001.txt
│       └── ...
└── classes.txt

将所有图像移至train/images,转换后的标签移至train/labels即可。

方案2:用YOLO官方工具直接下载转换

使用ultralytics库自带的工具,直接获取符合格式的Open Images数据:

  1. 安装依赖:
pip install ultralytics
  1. 执行下载命令(以Apple、Orange为例,限制100张训练集):
yolo data download dataset=open-images-v7 classes=Apple,Orange split=train limit=100

该命令会自动下载图像并生成YOLOv8兼容的标签,数据集结构直接满足训练要求。


内容的提问来源于stack exchange,提问作者OXYNO FIRING

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 22:07:33