You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何利用YOLOv4的ext_output参数返回的边界框坐标裁剪图像?

利用YOLOv4 ext_output的边界框坐标裁剪图像

首先明确ext_output返回的坐标定义:

  • left_x: 边界框左上角的x轴像素坐标
  • top_y: 边界框左上角的y轴像素坐标
  • width: 边界框的水平像素宽度
  • height: 边界框的垂直像素高度

要完成裁剪,需先计算边界框的右下角坐标:

right_x = left_x + width
bottom_y = top_y + height

步骤1:解析ext_output的结果

从输出文本中提取每个目标的坐标信息,用Python正则表达式可快速实现:

import re

# 示例ext_output输出文本
output_text = """dog: 99%    (left_x:   62   top_y:  261   width:  138   height:   89)
person: 99% (left_x:  191   top_y:   93   width:   82   height:  285)
horse: 79%  (left_x:  401   top_y:  131   width:  203   height:  207)"""

# 匹配每个目标的信息
pattern = r'(\w+): (\d+)%.*left_x:\s*(\d+)\s*top_y:\s*(\d+)\s*width:\s*(\d+)\s*height:\s*(\d+)'
matches = re.findall(pattern, output_text)

# 整理成字典列表
detections = []
for match in matches:
    cls, conf, left_x, top_y, width, height = match
    detections.append({
        'class': cls,
        'confidence': int(conf),
        'left_x': int(left_x),
        'top_y': int(top_y),
        'width': int(width),
        'height': int(height)
    })

步骤2:用OpenCV裁剪图像

通过数组切片直接实现裁剪:

import cv2

# 读取原始图像
img = cv2.imread('input_image.jpg')

for det in detections:
    left_x = det['left_x']
    top_y = det['top_y']
    right_x = left_x + det['width']
    bottom_y = top_y + det['height']
    
    # OpenCV图像格式为[行, 列],对应[y, x]
    cropped_img = img[top_y:bottom_y, left_x:right_x]
    
    # 保存裁剪后的图像
    cv2.imwrite(f"{det['class']}_{det['confidence']}.jpg", cropped_img)

步骤3:用PIL裁剪图像

PIL的裁剪区域参数为(left, top, right, bottom):

from PIL import Image

# 读取原始图像
img = Image.open('input_image.jpg')

for det in detections:
    left_x = det['left_x']
    top_y = det['top_y']
    right_x = left_x + det['width']
    bottom_y = top_y + det['height']
    
    # 执行裁剪
    cropped_img = img.crop((left_x, top_y, right_x, bottom_y))
    
    # 保存裁剪后的图像
    cropped_img.save(f"{det['class']}_{det['confidence']}.jpg")

注意事项

  • 裁剪前需检查right_x和bottom_y是否超出图像尺寸,避免报错:
    # OpenCV格式下获取图像尺寸
    img_height, img_width = img.shape[:2]
    right_x = min(right_x, img_width)
    bottom_y = min(bottom_y, img_height)
    
  • 确保坐标为整数,因为图像像素是离散的整数位置。

内容的提问来源于stack exchange,提问作者Chopper

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 01:35:22