使用Albumentations做YOLOv1图像增强时出现边界框坐标异常错误
我是深度学习与计算机视觉领域的新手,刚完成YOLOv1的实现教程,目前尝试在新数据集上应用该模型。编写代码处理图像与边界框后,确认train_ds2的标签无负值,但经过Albumentations增强得到train_ds3后,遍历数据集时出现如下错误:
InvalidArgumentError: {{function_node _wrapped__IteratorGetNext_output_types_2_device/job:localhost/replica:0/task:0/device:CPU:0}} ValueError: Expected y_min for bbox (0.203125, -0.002777785062789917, 0.8187500238418579, 0.8694444596767426, 5.0) to be in the range [0.0, 1.0], got -0.002777785062789917。
即使注释掉所有图像增强操作,问题仍然存在,想请教我忽略了什么问题?
相关代码
def get_bboxes(filename): name = filename[:-4] indices = meta_df[meta_df['new_img_id'] == float(name)] bounding_boxes = [] for index, row in indices.iterrows(): x = float(row['x']) y = float(row['y']) w = float(row['width']) h = float(row['height']) im_w = float(row['img_width']) im_h = float(row['img_height']) class_id = int(row['cat_id']) bounding_box = [(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h, class_id] bounding_boxes.append(bounding_box) return tf.convert_to_tensor(bounding_boxes, dtype=tf.float32) def generate_output(bounding_boxes): output_label = np.zeros([int(SPLIT_SIZE), int(SPLIT_SIZE), int(N_CLASSES+5)]) for b in range(len(bounding_boxes)): grid_x = bounding_boxes[...,b,0]*SPLIT_SIZE grid_y = bounding_boxes[...,b,1]*SPLIT_SIZE i = int(grid_x) j = int(grid_y) output_label[i, j, 0:5] = [1., grid_x%1, grid_y%1, bounding_boxes[...,b,2], bounding_boxes[...,b,3]] output_label[i, j, 5+int(bounding_boxes[...,b,4])] = 1. return tf.convert_to_tensor(output_label, tf.float32) def get_imboxes(im_path, map): img = tf.io.decode_jpeg(tf.io.read_file('./new_dataset/'+im_path)) img = tf.cast(tf.image.resize(img, [H,W]), dtype=tf.float32) bboxes = tf.numpy_function(func=get_bboxes, inp=[im_path], Tout=tf.float32) return img, bboxes train_ds2 = train_ds1.map(get_imboxes) val_ds2 = val_ds1.map(get_imboxes) transforms = A.Compose([ A.Resize(H,W), A.RandomCrop( width = np.random.randint(int(0.9*W), W), height = np.random.randint(int(0.9*H), H), p=0.5), A.RandomScale(scale_limit=0.1, interpolation=cv.INTER_LANCZOS4, p=0.5), A.HorizontalFlip(p=0.5), A.Resize(H,W) ], bbox_params =A.BboxParams(format='yolo')) def aug_albument(image, bboxes): augmented = transforms(image = image, bboxes = bboxes) return [tf.convert_to_tensor(augmented['image'], dtype=tf.float32), tf.convert_to_tensor(augmented['bboxes'], dtype=tf.float32)] def process_data(image, bboxes): aug = tf.numpy_function(func=aug_albument, inp=[image, bboxes], Tout=(tf.float32, tf.float32)) return aug[0], aug[1] train_ds3 = train_ds2.map(process_data)
1. Albumentations边界框格式解析错误
传入Albumentations的bbox包含5个元素(x_center, y_center, w, h, class_id),但设置format='yolo'时,Albumentations默认仅期望每个bbox是4个坐标值,class标签需要单独通过label_fields参数传递。未设置该参数会导致Albumentations错误解析bbox内容,将class_id误判为坐标值,进而出现坐标越界的错误。
解决代码调整:
# 1. 修改get_bboxes,拆分坐标与class标签 def get_bboxes(filename): name = filename[:-4] indices = meta_df[meta_df['new_img_id'] == float(name)] bbox_coords = [] class_ids = [] for index, row in indices.iterrows(): x = float(row['x']) y = float(row['y']) w = float(row['width']) h = float(row['height']) im_w = float(row['img_width']) im_h = float(row['img_height']) class_id = int(row['cat_id']) # 仅存储坐标部分 bbox_coord = [(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h] bbox_coords.append(bbox_coord) class_ids.append(class_id) return tf.convert_to_tensor(bbox_coords, dtype=tf.float32), tf.convert_to_tensor(class_ids, dtype=tf.int32) # 2. 修改get_imboxes,返回图像、坐标、标签 def get_imboxes(im_path, map): img = tf.io.decode_jpeg(tf.io.read_file('./new_dataset/'+im_path)) img = tf.cast(tf.image.resize(img, [H,W]), dtype=tf.float32) bboxes, class_ids = tf.numpy_function(func=get_bboxes, inp=[im_path], Tout=(tf.float32, tf.int32)) return img, bboxes, class_ids # 3. 更新数据集映射 train_ds2 = train_ds1.map(get_imboxes) val_ds2 = val_ds1.map(get_imboxes) # 4. 修改transforms,添加label_fields参数 transforms = A.Compose([ A.Resize(H,W), A.RandomCrop( width = np.random.randint(int(0.9*W), W), height = np.random.randint(int(0.9*H), H), p=0.5), A.RandomScale(scale_limit=0.1, interpolation=cv.INTER_LANCZOS4, p=0.5), A.HorizontalFlip(p=0.5), A.Resize(H,W) ], bbox_params =A.BboxParams(format='yolo', label_fields=['class_ids'])) # 5. 修改aug_albument,适配拆分后的输入 def aug_albument(image, bboxes, class_ids): # 转换为numpy数组处理 augmented = transforms(image = image.numpy(), bboxes = bboxes.numpy(), class_ids = class_ids.numpy()) # 重新组合坐标与class_id augmented_bboxes = [bbox + [cls] for bbox, cls in zip(augmented['bboxes'], augmented['class_ids'])] return ( tf.convert_to_tensor(augmented['image'], dtype=tf.float32), tf.convert_to_tensor(augmented_bboxes, dtype=tf.float32) ) # 6. 修改process_data,适配三个输入参数 def process_data(image, bboxes, class_ids): aug = tf.numpy_function(func=aug_albument, inp=[image, bboxes, class_ids], Tout=(tf.float32, tf.float32)) return aug[0], aug[1] train_ds3 = train_ds2.map(process_data)
2. 浮点精度误差导致的微小负值
即使bbox计算逻辑正确,浮点运算的精度误差可能导致原本应为0.0的值变成极小的负值(如错误中的-0.002777)。可以通过裁剪操作限制所有坐标值在[0.0, 1.0]范围内。
解决代码示例:
在get_bboxes中添加裁剪:
bbox_coord = np.clip([(x+w/2)/(im_w), (y+h/2)/(im_h), w/im_w, h/im_h], 0.0, 1.0)
或在aug_albument中处理增强后的bbox:
augmented_bboxes = np.clip(augmented_bboxes, 0.0, 1.0)
3. generate_output函数的索引错误
原代码中bounding_boxes[...,b,0]的索引方式有误,bounding_boxes的形状是(num_bboxes, 5),正确索引应为bounding_boxes[b, 0]。错误索引可能导致计算出异常的网格坐标,在增强后暴露问题。
修改后的generate_output:
def generate_output(bounding_boxes): output_label = np.zeros([int(SPLIT_SIZE), int(SPLIT_SIZE), int(N_CLASSES+5)]) for b in range(len(bounding_boxes)): # 修正索引方式 grid_x = bounding_boxes[b,0]*SPLIT_SIZE grid_y = bounding_boxes[b,1]*SPLIT_SIZE # 确保网格索引在合法范围内 i = max(0, min(int(grid_x), SPLIT_SIZE-1)) j = max(0, min(int(grid_y), SPLIT_SIZE-1)) output_label[i, j, 0:5] = [1., grid_x%1, grid_y%1, bounding_boxes[b,2], bounding_boxes[b,3]] output_label[i, j, 5+int(bounding_boxes[b,4])] = 1. return tf.convert_to_tensor(output_label, tf.float32)
内容的提问来源于stack exchange,提问作者Extra_Caterpillar

