如何基于xml标注的bounding box从单张图像中裁剪得到全部n张子图
问题根源
你的代码在遍历bounding box的循环里直接执行了return语句,函数第一次拿到裁剪图就会直接终止运行,因此只能返回第一个bbox对应的结果。另外你代码中定义的images_filenames、annotations_filenames两个变量没有被实际使用,可以直接删除。
修改方案
初始化一个空列表存储所有裁剪结果,遍历完全部bbox后统一返回列表即可,修改后的代码如下:
def get_all_cropped_images(image_file, images_path = "data/split/train/images", annotations_path = "data/split/train/annotations"): # 读取原图像 image = cv2.imread(os.path.join(images_path, image_file)) # 匹配对应xml标注文件,用rsplit避免文件名含多个点时匹配错误 annotation_filename = image_file.rsplit(sep='.', maxsplit=1)[0] + '.xml' # 读取标注信息 bboxes, labels = read_annotation_file(annotations_path, annotation_filename) # 初始化列表存储所有裁剪结果 cropped_images = [] for idx, label in enumerate(labels): try: cropped_img = imcrop(image, bboxes[idx]).copy() # 如果需要同时存储对应标签,可改为append( (cropped_img, label) ) cropped_images.append(cropped_img) except: # 若不需要严格报错,可注释掉raise,加日志打印后继续处理剩余bbox raise # 遍历完成后统一返回所有裁剪图 return cropped_images
使用说明
修改后函数返回值为包含所有n张裁剪图的列表,你可以直接遍历列表做后续OCR识别,也可以通过索引获取对应位置的裁剪结果。如果需要同时关联每个裁剪图对应的标注标签,可将append逻辑改为存入(裁剪图, 标签)的元组即可。
内容的提问来源于stack exchange,提问作者Gustavo Scholze
相关产品推荐
相关产品推荐

