You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将未知格式的人脸表情数据集标签转换为Yolov5标注格式

标注转换方案

1 YOLOv5标注格式要求

YOLOv5要求每张图片对应一个同名的.txt标注文件,统一存放在labels目录下,单条标注行的格式为:
类别ID 归一化中心点x坐标 归一化中心点y坐标 归一化框宽度 归一化框高度
所有坐标值均需要归一化到0~1范围内。

2 转换核心步骤

  • 提前获取所有待标注图片的宽、高参数,用于坐标归一化
  • 逐行解析原始标注,提取图片名、人脸框四点坐标、人脸框置信度、表情标签(作为YOLO的类别ID)
  • 把原始的左上右下绝对坐标转换为YOLO格式的中心+宽高绝对坐标,再除以图片宽高完成归一化
  • 过滤不合格标注:置信度低于阈值的标注、宽高为0的无效框、归一化后超出0~1范围的坐标都可以直接剔除
  • 按图片名分组,同一张图片的所有标注行写入对应同名的.txt文件

3 示例转换代码(Python)

import os
from PIL import Image

# 路径配置,需自行替换为你的实际路径
IMG_DIR = "存放原始图片的目录路径"
RAW_LABEL_PATH = "原始标注文件的完整路径"
OUT_LABEL_DIR = "输出YOLO格式标注的目录,请提前创建"
# 人脸框置信度阈值,根据你的原始标注置信度范围调整,示例中置信度为22.9362,可自行设定合适阈值
CONF_THRESH = 10.0 

if not os.path.exists(OUT_LABEL_DIR):
    os.makedirs(OUT_LABEL_DIR)

# 逐行处理原始标注
with open(RAW_LABEL_PATH, 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        parts = line.split()
        img_name = parts[0]
        box_top = int(parts[2])
        box_left = int(parts[3])
        box_right = int(parts[4])
        box_bottom = int(parts[5])
        conf = float(parts[6])
        cls_id = int(parts[7])

        # 置信度过滤
        if conf < CONF_THRESH:
            continue
        
        # 读取对应图片的宽高
        img_path = os.path.join(IMG_DIR, img_name)
        if not os.path.exists(img_path):
            print(f"未找到图片{img_name},跳过该标注")
            continue
        try:
            with Image.open(img_path) as img:
                img_w, img_h = img.size
        except Exception as e:
            print(f"读取图片{img_name}失败:{e},跳过该标注")
            continue
        
        # 转换为中心宽高坐标
        abs_x_center = (box_left + box_right) / 2
        abs_y_center = (box_top + box_bottom) / 2
        abs_w = box_right - box_left
        abs_h = box_bottom - box_top

        # 坐标归一化
        x_center = abs_x_center / img_w
        y_center = abs_y_center / img_h
        w = abs_w / img_w
        h = abs_h / img_h

        # 合法性校验,避免越界
        x_center = max(0.0, min(1.0, x_center))
        y_center = max(0.0, min(1.0, y_center))
        w = max(0.0, min(1.0, w))
        h = max(0.0, min(1.0, h))
        if w <= 1e-6 or h <= 1e-6:
            continue
        
        # 写入标注文件
        txt_filename = os.path.splitext(img_name)[0] + '.txt'
        out_txt_path = os.path.join(OUT_LABEL_DIR, txt_filename)
        with open(out_txt_path, 'a', encoding='utf-8') as out_f:
            out_f.write(f"{cls_id} {x_center:.6f} {y_center:.6f} {w:.6f} {h:.6f}\n")

4 转换结果示例

以你给出的标注示例angry_actor_104.jpg 0 28 113 226 141 22.9362 0为例,假设该图片分辨率为400×300:

  • 计算得到绝对中心点x=(113+226)/2=169.5,归一化后为169.5/400=0.42375
  • 绝对中心点y=(28+141)/2=84.5,归一化后为84.5/300≈0.281667
  • 框宽度=226-113=113,归一化后为113/400=0.2825
  • 框高度=141-28=113,归一化后为113/300≈0.376667
    最终写入angry_actor_104.txt的标注内容为:
    0 0.423750 0.281667 0.282500 0.376667

内容的提问来源于stack exchange,提问作者Phil Harmony

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 17:15:03