如何实现OpenCV+YOLO目标检测的单次语音播报反馈功能
目标检测语音反馈重复播报问题解决
问题说明
基于OpenCV、YOLO框架与pyttsx3开发的带语音反馈目标检测系统,当前存在识别到物体(如person)时持续重复播报的问题。需求为:每个识别到的物体仅播报一次,新物体(如bottle)进入画面时也仅播报一次。
解决思路
- 维护一个集合记录已播报的物体类别,避免同一类别重复播报
- 将pyttsx3引擎初始化移到循环外部,减少资源消耗
- 每次检测到物体后,先判断该类别是否已在播报记录中,未记录则执行播报并加入集合
- (可选)若需要支持物体离开画面后重新进入再次播报,可添加物体消失检测逻辑,将对应类别从记录集合中移除
修改后的完整代码
import cv2 import numpy as np import pyttsx3 net = cv2.dnn.readNet('yolov3-tiny.weights', 'yolov3-tiny.cfg') classes = [] with open("coco.names.txt", "r") as f: classes = f.read().splitlines() cap = cv2.VideoCapture(0) font = cv2.FONT_HERSHEY_PLAIN colors = np.random.uniform(0, 255, size=(100, 3)) # 初始化语音引擎(移到循环外) engine = pyttsx3.init() # 记录已播报的物体类别 announced_classes = set() while True: _, img = cap.read() height, width, _ = img.shape blob = cv2.dnn.blobFromImage(img, 1/255, (416, 416), (0,0,0), swapRB=True, crop=False) net.setInput(blob) output_layers_names = net.getUnconnectedOutLayersNames() layerOutputs = net.forward(output_layers_names) boxes = [] confidences = [] class_ids = [] for output in layerOutputs: for detection in output: scores = detection[5:] class_id = np.argmax(scores) confidence = scores[class_id] if confidence > 0.2: center_x = int(detection[0]*width) center_y = int(detection[1]*height) w = int(detection[2]*width) h = int(detection[3]*height) x = int(center_x - w/2) y = int(center_y - h/2) boxes.append([x, y, w, h]) confidences.append((float(confidence))) class_ids.append(class_id) indexes = cv2.dnn.NMSBoxes(boxes, confidences, 0.2, 0.4) # 记录当前帧检测到的类别 current_detected_classes = set() if len(indexes)>0: for i in indexes.flatten(): x, y, w, h = boxes[i] label = str(classes[class_ids[i]]) current_detected_classes.add(label) confidence = str(round(confidences[i],2)) color = colors[i] cv2.rectangle(img, (x,y), (x+w, y+h), color, 2) cv2.putText(img, label + " " + confidence, (x, y+20), font, 2, (255,255,255), 2) # 遍历当前检测到的类别,仅播报未播报过的 for cls in current_detected_classes: if cls not in announced_classes: engine.say(f"检测到{cls}") engine.runAndWait() announced_classes.add(cls) # 可选:如果需要物体消失后重新播报,取消以下注释 # for cls in list(announced_classes): # if cls not in current_detected_classes: # announced_classes.remove(cls) cv2.imshow('Image', img) key = cv2.waitKey(1) if key==27: break cap.release() cv2.destroyAllWindows()
关键修改点说明
- 语音引擎优化:将
engine = pyttsx3.init()移到循环外部,避免每次循环创建新引擎,提升性能同时减少播报延迟 - 播报记录集合:新增
announced_classes集合存储已播报的类别,确保同一类别仅触发一次播报 - 当前检测类别统一处理:在循环内收集当前帧所有检测到的类别,统一执行播报逻辑,避免同一帧内同一类别多次播报
- 可选消失检测逻辑:注释部分提供了物体离开画面后,将对应类别从记录中移除的逻辑,取消注释即可支持物体重新进入时再次播报
内容的提问来源于stack exchange,提问作者Aya Mohammed
相关产品推荐
相关产品推荐

