You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更新裁剪后图像的YOLOv5标注.txt文件以适配二级模型训练?

批量更新YOLOv5裁剪后图像的标注坐标方案

核心逻辑

利用原图的YOLOv5标注信息,结合里程表裁剪区域的像素位置和尺寸,批量转换数字的归一化坐标到裁剪后的图像坐标系,同时自动添加整个裁剪区域的标注条目,无需手动重新标注2000张图。

具体操作步骤

1. 获取里程表的裁剪区域坐标

如果原图标注中已经包含里程表的类别(比如类ID为0),直接从原图的.txt标注文件中解析出它的绝对像素坐标(左上角(x1,y1)、右下角(x2,y2)):

  • 先读取原图的宽高(img_w, img_h)
  • 把标注里的归一化坐标转成绝对坐标:
    # 从标注行解析
    cls_id, norm_x, norm_y, norm_w, norm_h = map(float, line.split())
    x_center = norm_x * img_w
    y_center = norm_y * img_h
    w = norm_w * img_w
    h = norm_h * img_h
    x1 = x_center - w/2
    y1 = y_center - h/2
    x2 = x_center + w/2
    y2 = y_center + h/2
    
    得到的(x1,y1,x2,y2)就是裁剪里程表的区域范围。

2. 转换数字的标注坐标

裁剪后的图像尺寸为crop_w = x2 - x1、crop_h = y2 - y1,对原图中每个数字的标注做如下转换:

  • 先将数字的归一化坐标转为原图上的绝对坐标:num_x_center = norm_num_x * img_w、num_y_center = norm_num_y * img_h、num_w = norm_num_w * img_w、num_h = norm_num_h * img_h
  • 计算数字在裁剪图中的绝对中心坐标:crop_num_x = num_x_center - x1、crop_num_y = num_y_center - y1
  • 转为裁剪图的归一化坐标:
    new_norm_x = crop_num_x / crop_w
    new_norm_y = crop_num_y / crop_h
    new_norm_w = num_w / crop_w
    new_norm_h = num_h / crop_h
    
    可选:如果数字区域部分超出裁剪范围,直接过滤这类标注(裁剪图中看不到的目标无需保留)。

3. 添加裁剪区域的标注条目

裁剪后的整张图就是里程表区域,所以对应的YOLOv5标注为:

  • 假设里程表的类别ID设为0,数字为1(可根据你的分类需求调整)
  • 标注行格式:0 0.5 0.5 1.0 1.0(中心在图像正中间,宽高占满整个裁剪图)

4. 批量处理脚本

直接用以下Python脚本批量处理所有图像和标注:

import os
from PIL import Image

# 配置参数,根据你的路径和类别ID修改
ORIG_IMG_DIR = "原始图像文件夹路径"
ORIG_LABEL_DIR = "原始标注文件夹路径"
CROP_IMG_DIR = "裁剪后图像保存路径"
CROP_LABEL_DIR = "裁剪后标注保存路径"
ODOMETER_CLAS_ID = 0  # 原图中里程表的类别ID
NUM_CLASS_ID = 1       # 裁剪图中数字的类别ID
NEW_ODOMETER_ID = 0    # 裁剪图中里程表区域的类别ID

# 创建输出文件夹
os.makedirs(CROP_IMG_DIR, exist_ok=True)
os.makedirs(CROP_LABEL_DIR, exist_ok=True)

# 遍历所有图像
for img_filename in os.listdir(ORIG_IMG_DIR):
    if not img_filename.lower().endswith((".jpg", ".png", ".jpeg")):
        continue
    
    img_path = os.path.join(ORIG_IMG_DIR, img_filename)
    label_path = os.path.join(ORIG_LABEL_DIR, img_filename.rsplit('.', 1)[0] + ".txt")
    
    if not os.path.exists(label_path):
        continue
    
    # 读取原图尺寸
    img = Image.open(img_path)
    img_w, img_h = img.size
    
    odometer_box = None
    num_labels = []
    
    # 解析原图标注
    with open(label_path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if not line:
                continue
            parts = list(map(float, line.split()))
            cls_id = int(parts[0])
            if cls_id == ODOMETER_CLAS_ID:
                # 转换为绝对裁剪坐标
                norm_x, norm_y, norm_w, norm_h = parts[1:]
                x_center = norm_x * img_w
                y_center = norm_y * img_h
                w = norm_w * img_w
                h = norm_h * img_h
                odometer_box = (x_center - w/2, y_center - h/2, x_center + w/2, y_center + h/2)
            else:
                # 收集数字标注
                num_labels.append(parts)
    
    if not odometer_box:
        continue
    
    # 裁剪图像
    x1, y1, x2, y2 = map(int, odometer_box)
    crop_img = img.crop((x1, y1, x2, y2))
    crop_w = x2 - x1
    crop_h = y2 - y1
    crop_img.save(os.path.join(CROP_IMG_DIR, img_filename))
    
    # 生成裁剪后的标注
    crop_label_content = []
    # 添加里程表区域标注
    crop_label_content.append(f"{NEW_ODOMETER_ID} 0.5 0.5 1.0 1.0")
    # 转换数字标注
    for label in num_labels:
        cls_id, norm_x, norm_y, norm_w, norm_h = label
        # 转为原图绝对坐标
        num_x_center = norm_x * img_w
        num_y_center = norm_y * img_h
        num_w = norm_w * img_w
        num_h = norm_h * img_h
        # 计算裁剪图中的坐标
        crop_x = num_x_center - x1
        crop_y = num_y_center - y1
        # 检查数字是否完全在裁剪区域内(可选)
        if (crop_x - num_w/2 >= 0 and crop_x + num_w/2 <= crop_w and
            crop_y - num_h/2 >= 0 and crop_y + num_h/2 <= crop_h):
            new_norm_x = crop_x / crop_w
            new_norm_y = crop_y / crop_h
            new_norm_w = num_w / crop_w
            new_norm_h = num_h / crop_h
            # 保留6位小数,符合YOLO格式
            crop_label_content.append(f"{NUM_CLASS_ID} {new_norm_x:.6f} {new_norm_y:.6f} {new_norm_w:.6f} {new_norm_h:.6f}")
    
    # 保存标注文件
    with open(os.path.join(CROP_LABEL_DIR, img_filename.rsplit('.', 1)[0] + ".txt"), "w", encoding="utf-8") as f:
        f.write("\n".join(crop_label_content))

5. 验证结果

随机挑选几张裁剪后的图像和对应的标注,用YOLOv5自带的可视化工具(比如utils/plots.py里的plot_one_box函数)检查标注位置是否准确,确保坐标转换无误。

内容的提问来源于stack exchange,提问作者Sid_J1996

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 03:05:19