如何利用YOLOv4的ext_output参数返回的边界框坐标裁剪图像?
利用YOLOv4 ext_output的边界框坐标裁剪图像
首先明确ext_output返回的坐标定义:
left_x: 边界框左上角的x轴像素坐标top_y: 边界框左上角的y轴像素坐标width: 边界框的水平像素宽度height: 边界框的垂直像素高度
要完成裁剪,需先计算边界框的右下角坐标:
right_x = left_x + width bottom_y = top_y + height
步骤1:解析ext_output的结果
从输出文本中提取每个目标的坐标信息,用Python正则表达式可快速实现:
import re # 示例ext_output输出文本 output_text = """dog: 99% (left_x: 62 top_y: 261 width: 138 height: 89) person: 99% (left_x: 191 top_y: 93 width: 82 height: 285) horse: 79% (left_x: 401 top_y: 131 width: 203 height: 207)""" # 匹配每个目标的信息 pattern = r'(\w+): (\d+)%.*left_x:\s*(\d+)\s*top_y:\s*(\d+)\s*width:\s*(\d+)\s*height:\s*(\d+)' matches = re.findall(pattern, output_text) # 整理成字典列表 detections = [] for match in matches: cls, conf, left_x, top_y, width, height = match detections.append({ 'class': cls, 'confidence': int(conf), 'left_x': int(left_x), 'top_y': int(top_y), 'width': int(width), 'height': int(height) })
步骤2:用OpenCV裁剪图像
通过数组切片直接实现裁剪:
import cv2 # 读取原始图像 img = cv2.imread('input_image.jpg') for det in detections: left_x = det['left_x'] top_y = det['top_y'] right_x = left_x + det['width'] bottom_y = top_y + det['height'] # OpenCV图像格式为[行, 列],对应[y, x] cropped_img = img[top_y:bottom_y, left_x:right_x] # 保存裁剪后的图像 cv2.imwrite(f"{det['class']}_{det['confidence']}.jpg", cropped_img)
步骤3:用PIL裁剪图像
PIL的裁剪区域参数为(left, top, right, bottom):
from PIL import Image # 读取原始图像 img = Image.open('input_image.jpg') for det in detections: left_x = det['left_x'] top_y = det['top_y'] right_x = left_x + det['width'] bottom_y = top_y + det['height'] # 执行裁剪 cropped_img = img.crop((left_x, top_y, right_x, bottom_y)) # 保存裁剪后的图像 cropped_img.save(f"{det['class']}_{det['confidence']}.jpg")
注意事项
- 裁剪前需检查
right_x和bottom_y是否超出图像尺寸,避免报错:# OpenCV格式下获取图像尺寸 img_height, img_width = img.shape[:2] right_x = min(right_x, img_width) bottom_y = min(bottom_y, img_height) - 确保坐标为整数,因为图像像素是离散的整数位置。
内容的提问来源于stack exchange,提问作者Chopper
相关产品推荐
相关产品推荐

