You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现OpenCV+YOLO目标检测的单次语音播报反馈功能

目标检测语音反馈重复播报问题解决

问题说明

基于OpenCV、YOLO框架与pyttsx3开发的带语音反馈目标检测系统,当前存在识别到物体(如person)时持续重复播报的问题。需求为:每个识别到的物体仅播报一次,新物体(如bottle)进入画面时也仅播报一次。

解决思路

  • 维护一个集合记录已播报的物体类别,避免同一类别重复播报
  • 将pyttsx3引擎初始化移到循环外部,减少资源消耗
  • 每次检测到物体后,先判断该类别是否已在播报记录中,未记录则执行播报并加入集合
  • (可选)若需要支持物体离开画面后重新进入再次播报,可添加物体消失检测逻辑,将对应类别从记录集合中移除

修改后的完整代码

import cv2
import numpy as np
import pyttsx3

net = cv2.dnn.readNet('yolov3-tiny.weights', 'yolov3-tiny.cfg')

classes = []
with open("coco.names.txt", "r") as f:
   classes = f.read().splitlines()

cap = cv2.VideoCapture(0)
font = cv2.FONT_HERSHEY_PLAIN
colors = np.random.uniform(0, 255, size=(100, 3))

# 初始化语音引擎(移到循环外)
engine = pyttsx3.init()
# 记录已播报的物体类别
announced_classes = set()

while True:
    _, img = cap.read()
    height, width, _ = img.shape

    blob = cv2.dnn.blobFromImage(img, 1/255, (416, 416), (0,0,0), swapRB=True, crop=False)
    net.setInput(blob)
    output_layers_names = net.getUnconnectedOutLayersNames()
    layerOutputs = net.forward(output_layers_names)

    boxes = []
    confidences = []
    class_ids = []

    for output in layerOutputs:
        for detection in output:
            scores = detection[5:]
            class_id = np.argmax(scores)
            confidence = scores[class_id]
            if confidence > 0.2:
                center_x = int(detection[0]*width)
                center_y = int(detection[1]*height)
                w = int(detection[2]*width)
                h = int(detection[3]*height)

                x = int(center_x - w/2)
                y = int(center_y - h/2)

                boxes.append([x, y, w, h])
                confidences.append((float(confidence)))
                class_ids.append(class_id)

    indexes = cv2.dnn.NMSBoxes(boxes, confidences, 0.2, 0.4)
    # 记录当前帧检测到的类别
    current_detected_classes = set()

    if len(indexes)>0:
        for i in indexes.flatten():
            x, y, w, h = boxes[i]
            label = str(classes[class_ids[i]])
            current_detected_classes.add(label)
            confidence = str(round(confidences[i],2))
            color = colors[i]
            cv2.rectangle(img, (x,y), (x+w, y+h), color, 2)
            cv2.putText(img, label + " " + confidence, (x, y+20), font, 2, (255,255,255), 2)

        # 遍历当前检测到的类别,仅播报未播报过的
        for cls in current_detected_classes:
            if cls not in announced_classes:
                engine.say(f"检测到{cls}")
                engine.runAndWait()
                announced_classes.add(cls)

    # 可选:如果需要物体消失后重新播报,取消以下注释
    # for cls in list(announced_classes):
    #     if cls not in current_detected_classes:
    #         announced_classes.remove(cls)

    cv2.imshow('Image', img)
    key = cv2.waitKey(1)
    if key==27:
        break

cap.release()
cv2.destroyAllWindows()

关键修改点说明

  • 语音引擎优化:将engine = pyttsx3.init()移到循环外部,避免每次循环创建新引擎,提升性能同时减少播报延迟
  • 播报记录集合:新增announced_classes集合存储已播报的类别,确保同一类别仅触发一次播报
  • 当前检测类别统一处理:在循环内收集当前帧所有检测到的类别,统一执行播报逻辑,避免同一帧内同一类别多次播报
  • 可选消失检测逻辑:注释部分提供了物体离开画面后,将对应类别从记录中移除的逻辑,取消注释即可支持物体重新进入时再次播报

内容的提问来源于stack exchange,提问作者Aya Mohammed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 00:10:29