You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用YOLOv5输出的xmin/ymin/xmax/ymax在截屏画面绘制边界框

解决方法

你可以直接通过OpenCV内置的rectangle方法使用获取到的xmin/ymin/xmax/ymax坐标绘制 bounding box,以下是修改后的完整可运行代码:

import torch
import numpy as np
import threading
import time
import cv2
from mss import mss

# 全局参数
bounding_box = {'top': 340, 'left': 650, 'width': 350, 'height': 400}
# 初始化全局共享变量
scr_img = np.zeros((bounding_box['height'], bounding_box['width'], 3), np.uint8)
detection_boxes = []

# 加载YOLOv5模型
model = torch.hub.load('ultralytics/yolov5', 'yolov5s', pretrained=True).autoshape()
sct = mss()

class ScreenGrab(threading.Thread):
    def __init__(self, id_manager):
        threading.Thread.__init__(self)
        self.id_manager = int(id_manager)

    def run(self):
        global scr_img, detection_boxes
        while True:
            sct_img = sct.grab(bounding_box)
            # 去掉mss返回的alpha通道,保留BGR三通道
            scr_img = np.array(sct_img)[:, :, :3]
            
            # 遍历检测结果绘制框和文本
            for box in detection_boxes:
                xmin, ymin, xmax, ymax, conf = box
                # 绘制 bounding box,绿色,线宽2
                cv2.rectangle(scr_img, (int(xmin), int(ymin)), (int(xmax), int(ymax)), (0, 255, 0), 2)
                # 绘制类别和置信度文本
                cv2.putText(scr_img, f'person {conf:.2f}', (int(xmin), int(ymin)-10), 
                            cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0, 255, 0), 2)
            
            cv2.imshow('Testing', scr_img)
            if (cv2.waitKey(1) & 0xFF) == ord('q'):
                cv2.destroyAllWindows()
                break


class DetectionModule(threading.Thread):
    def __init__(self, id_manager):
        threading.Thread.__init__(self)
        self.id_manager = int(id_manager)

    def run(self):
        global scr_img, detection_boxes
        while True:
            # 推理检测
            results = model(scr_img)
            boxes = results.pandas().xyxy[0]
            # 筛选person类别
            person = boxes.loc[boxes['class'] == 0]
            # 更新检测框列表
            temp_boxes = []
            for _, row in person.iterrows():
                temp_boxes.append([row['xmin'], row['ymin'], row['xmax'], row['ymax'], row['confidence']])
            detection_boxes = temp_boxes


def ScreenGrabMain():
    thread = ScreenGrab(0)
    thread.start()
    time.sleep(0.5)

def DetectionMain2():
    thread = DetectionModule(0)
    thread.start()
    time.sleep(0.5)

if __name__ == "__main__":
    ScreenGrabMain()
    DetectionMain2()

关键修改说明

  • 新增全局共享变量detection_boxes存储检测到的行人框坐标和置信度,避免两个线程读写冲突
  • 对mss抓取的BGRA格式截图做通道裁剪,仅保留BGR三通道,匹配OpenCV的显示格式
  • 截屏线程在显示画面之前,遍历检测框列表,调用cv2.rectangle绘制边界框,调用cv2.putText标注类别和置信度
  • 检测线程推理完成后,筛选出行人目标的坐标并更新到detection_boxes列表中
  • 初始化了空白的scr_img变量,避免检测线程启动时读取空变量报错

如果需要调整框的颜色、线宽、文字大小,直接修改cv2.rectangle和cv2.putText的对应参数即可。

内容的提问来源于stack exchange,提问作者terry5546

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 12:57:04