You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在OpenCV+YOLOv3中实现新增检测物体的单次播报

解决重复播报已识别物体的问题

要实现仅在新物体出现时播报一次,核心是记录已播报过的物体,并在每帧检测时对比当前识别结果,只对新增的物体进行语音输出。具体修改步骤如下:

  1. 初始化已播报物体集合:用集合存储已播报的物体类别,自动避免重复项。
  2. 提前初始化语音引擎:将pyttsx3.init()移到循环外,避免每次帧都重新初始化引擎导致卡顿。
  3. 收集当前帧的所有识别物体:在绘制检测框之前,先提取当前帧所有唯一的识别类别。
  4. 对比并播报新物体:找出当前帧中未播报过的物体,进行语音播报并加入已播报集合。

修改后的完整代码

import cv2
import numpy as np
import pyttsx3

net = cv2.dnn.readNet('yolov3-tiny.weights', 'yolov3-tiny.cfg')

classes = []
with open("coco.names.txt", "r") as f:
    classes = f.read().splitlines()

cap = cv2.VideoCapture(0)
font = cv2.FONT_HERSHEY_PLAIN
colors = np.random.uniform(0, 255, size=(100, 3))

# 初始化已播报物体集合
announced_objects = set()
# 提前初始化语音引擎
engine = pyttsx3.init()

while True:
    _, img = cap.read()
    height, width, _ = img.shape

    blob = cv2.dnn.blobFromImage(img, 1/255, (416, 416), (0,0,0), swapRB=True, crop=False)
    net.setInput(blob)
    output_layers_names = net.getUnconnectedOutLayersNames()
    layerOutputs = net.forward(output_layers_names)

    boxes = []
    confidences = []
    class_ids = []

    for output in layerOutputs:
        for detection in output:
            scores = detection[5:]
            class_id = np.argmax(scores)
            confidence = scores[class_id]
            if confidence > 0.2:
                center_x = int(detection[0]*width)
                center_y = int(detection[1]*height)
                w = int(detection[2]*width)
                h = int(detection[3]*height)

                x = int(center_x - w/2)
                y = int(center_y - h/2)

                boxes.append([x, y, w, h])
                confidences.append((float(confidence)))
                class_ids.append(class_id)

    indexes = cv2.dnn.NMSBoxes(boxes, confidences, 0.2, 0.4)

    current_labels = set()
    if len(indexes) > 0:
        for i in indexes.flatten():
            x, y, w, h = boxes[i]
            label = str(classes[class_ids[i]])
            confidence = str(round(confidences[i], 2))
            color = colors[i]
            cv2.rectangle(img, (x,y), (x+w, y+h), color, 2)
            cv2.putText(img, label + " " + confidence, (x, y+20), font, 2, (255,255,255), 2)
            current_labels.add(label)
        
        # 找出新增的物体并播报
        new_objects = current_labels - announced_objects
        for obj in new_objects:
            engine.say(f"检测到{obj}")
            engine.runAndWait()
            announced_objects.add(obj)

    cv2.imshow('Image', img)
    key = cv2.waitKey(1)
    if key == 27:
        break

cap.release()
cv2.destroyAllWindows()

额外说明

  • 如果需要在物体消失后再次出现时重新播报,可以添加逻辑定期清空announced_objects集合,或者跟踪物体的存在状态(比如连续N帧未检测到则从集合中移除)。
  • 调整confidence阈值(当前为0.2)可以平衡检测精度和播报灵敏度,阈值越高误报越少,但可能漏检部分物体。

内容的提问来源于stack exchange,提问作者Aya Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 00:35:21