You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Albumentations做YOLOv1图像增强时出现边界框坐标异常错误

问题:YOLOv1数据集增强后出现边界框坐标负值错误

我是深度学习与计算机视觉领域的新手,刚完成YOLOv1的实现教程,目前尝试在新数据集上应用该模型。编写代码处理图像与边界框后,确认train_ds2的标签无负值,但经过Albumentations增强得到train_ds3后,遍历数据集时出现如下错误:

InvalidArgumentError: {{function_node _wrapped__IteratorGetNext_output_types_2_device/job:localhost/replica:0/task:0/device:CPU:0}} ValueError: Expected y_min for bbox (0.203125, -0.002777785062789917, 0.8187500238418579, 0.8694444596767426, 5.0) to be in the range [0.0, 1.0], got -0.002777785062789917。

即使注释掉所有图像增强操作,问题仍然存在,想请教我忽略了什么问题?

相关代码

def get_bboxes(filename):
  name = filename[:-4]
  indices = meta_df[meta_df['new_img_id'] == float(name)]
  bounding_boxes = []
  for index, row in indices.iterrows():
    x = float(row['x'])
    y = float(row['y'])
    w = float(row['width'])
    h = float(row['height'])
    im_w = float(row['img_width'])
    im_h = float(row['img_height'])
    class_id = int(row['cat_id'])
    bounding_box = [(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h, class_id]
    bounding_boxes.append(bounding_box)
  return tf.convert_to_tensor(bounding_boxes, dtype=tf.float32)

def generate_output(bounding_boxes):
  output_label = np.zeros([int(SPLIT_SIZE), int(SPLIT_SIZE), int(N_CLASSES+5)])
  for b in range(len(bounding_boxes)):
    grid_x = bounding_boxes[...,b,0]*SPLIT_SIZE
    grid_y = bounding_boxes[...,b,1]*SPLIT_SIZE

    i = int(grid_x)
    j = int(grid_y)

    output_label[i, j, 0:5] = [1., grid_x%1, grid_y%1, bounding_boxes[...,b,2], bounding_boxes[...,b,3]]
    output_label[i, j, 5+int(bounding_boxes[...,b,4])] = 1.

  return tf.convert_to_tensor(output_label, tf.float32)

def get_imboxes(im_path, map):
  img = tf.io.decode_jpeg(tf.io.read_file('./new_dataset/'+im_path))
  img = tf.cast(tf.image.resize(img, [H,W]), dtype=tf.float32)

  bboxes = tf.numpy_function(func=get_bboxes, inp=[im_path], Tout=tf.float32)

  return img, bboxes

train_ds2 = train_ds1.map(get_imboxes)
val_ds2 = val_ds1.map(get_imboxes)

transforms = A.Compose([
    A.Resize(H,W),
    A.RandomCrop(
        width = np.random.randint(int(0.9*W), W),
        height = np.random.randint(int(0.9*H), H), p=0.5),
    A.RandomScale(scale_limit=0.1, interpolation=cv.INTER_LANCZOS4, p=0.5),
    A.HorizontalFlip(p=0.5),
    A.Resize(H,W)
], bbox_params =A.BboxParams(format='yolo'))

def aug_albument(image, bboxes):
  augmented = transforms(image = image, bboxes = bboxes)
  return [tf.convert_to_tensor(augmented['image'], dtype=tf.float32), tf.convert_to_tensor(augmented['bboxes'], dtype=tf.float32)]

def process_data(image, bboxes):
  aug = tf.numpy_function(func=aug_albument, inp=[image, bboxes], Tout=(tf.float32, tf.float32))
  return aug[0], aug[1]

train_ds3 = train_ds2.map(process_data)
问题分析与解决

1. Albumentations边界框格式解析错误

传入Albumentations的bbox包含5个元素(x_center, y_center, w, h, class_id),但设置format='yolo'时,Albumentations默认仅期望每个bbox是4个坐标值,class标签需要单独通过label_fields参数传递。未设置该参数会导致Albumentations错误解析bbox内容,将class_id误判为坐标值,进而出现坐标越界的错误。

解决代码调整:

# 1. 修改get_bboxes,拆分坐标与class标签
def get_bboxes(filename):
  name = filename[:-4]
  indices = meta_df[meta_df['new_img_id'] == float(name)]
  bbox_coords = []
  class_ids = []
  for index, row in indices.iterrows():
    x = float(row['x'])
    y = float(row['y'])
    w = float(row['width'])
    h = float(row['height'])
    im_w = float(row['img_width'])
    im_h = float(row['img_height'])
    class_id = int(row['cat_id'])
    # 仅存储坐标部分
    bbox_coord = [(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h]
    bbox_coords.append(bbox_coord)
    class_ids.append(class_id)
  return tf.convert_to_tensor(bbox_coords, dtype=tf.float32), tf.convert_to_tensor(class_ids, dtype=tf.int32)

# 2. 修改get_imboxes,返回图像、坐标、标签
def get_imboxes(im_path, map):
  img = tf.io.decode_jpeg(tf.io.read_file('./new_dataset/'+im_path))
  img = tf.cast(tf.image.resize(img, [H,W]), dtype=tf.float32)
  bboxes, class_ids = tf.numpy_function(func=get_bboxes, inp=[im_path], Tout=(tf.float32, tf.int32))
  return img, bboxes, class_ids

# 3. 更新数据集映射
train_ds2 = train_ds1.map(get_imboxes)
val_ds2 = val_ds1.map(get_imboxes)

# 4. 修改transforms,添加label_fields参数
transforms = A.Compose([
    A.Resize(H,W),
    A.RandomCrop(
        width = np.random.randint(int(0.9*W), W),
        height = np.random.randint(int(0.9*H), H), p=0.5),
    A.RandomScale(scale_limit=0.1, interpolation=cv.INTER_LANCZOS4, p=0.5),
    A.HorizontalFlip(p=0.5),
    A.Resize(H,W)
], bbox_params =A.BboxParams(format='yolo', label_fields=['class_ids']))

# 5. 修改aug_albument,适配拆分后的输入
def aug_albument(image, bboxes, class_ids):
  # 转换为numpy数组处理
  augmented = transforms(image = image.numpy(), bboxes = bboxes.numpy(), class_ids = class_ids.numpy())
  # 重新组合坐标与class_id
  augmented_bboxes = [bbox + [cls] for bbox, cls in zip(augmented['bboxes'], augmented['class_ids'])]
  return (
      tf.convert_to_tensor(augmented['image'], dtype=tf.float32),
      tf.convert_to_tensor(augmented_bboxes, dtype=tf.float32)
  )

# 6. 修改process_data,适配三个输入参数
def process_data(image, bboxes, class_ids):
  aug = tf.numpy_function(func=aug_albument, inp=[image, bboxes, class_ids], Tout=(tf.float32, tf.float32))
  return aug[0], aug[1]

train_ds3 = train_ds2.map(process_data)

2. 浮点精度误差导致的微小负值

即使bbox计算逻辑正确,浮点运算的精度误差可能导致原本应为0.0的值变成极小的负值(如错误中的-0.002777)。可以通过裁剪操作限制所有坐标值在[0.0, 1.0]范围内。

解决代码示例:

在get_bboxes中添加裁剪:

bbox_coord = np.clip([(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h], 0.0, 1.0)

或在aug_albument中处理增强后的bbox:

augmented_bboxes = np.clip(augmented_bboxes, 0.0, 1.0)

3. generate_output函数的索引错误

原代码中bounding_boxes[...,b,0]的索引方式有误,bounding_boxes的形状是(num_bboxes, 5),正确索引应为bounding_boxes[b, 0]。错误索引可能导致计算出异常的网格坐标,在增强后暴露问题。

修改后的generate_output:

def generate_output(bounding_boxes):
  output_label = np.zeros([int(SPLIT_SIZE), int(SPLIT_SIZE), int(N_CLASSES+5)])
  for b in range(len(bounding_boxes)):
    # 修正索引方式
    grid_x = bounding_boxes[b,0]*SPLIT_SIZE
    grid_y = bounding_boxes[b,1]*SPLIT_SIZE

    # 确保网格索引在合法范围内
    i = max(0, min(int(grid_x), SPLIT_SIZE-1))
    j = max(0, min(int(grid_y), SPLIT_SIZE-1))

    output_label[i, j, 0:5] = [1., grid_x%1, grid_y%1, bounding_boxes[b,2], bounding_boxes[b,3]]
    output_label[i, j, 5+int(bounding_boxes[b,4])] = 1.

  return tf.convert_to_tensor(output_label, tf.float32)

内容的提问来源于stack exchange,提问作者Extra_Caterpillar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 15:45:55