如何修改Python面部情绪识别代码 过滤孤立预测以平滑输出结果
面部情绪识别结果平滑实现方案
核心思路
针对预测结果跳变的问题,我们通过连续判定逻辑过滤孤立的异常预测:
- 设定平滑阈值
n,代表需要连续命中同一情绪的次数 - 维护连续相同预测的计数,和当前稳定展示的情绪值
- 仅当连续
n次预测结果完全一致时,才更新界面展示的情绪内容,否则保留上一次稳定的展示结果
修改步骤
无需修改模型训练相关代码,仅替换display模式下的代码即可:
# emotions will be displayed on your face from the webcam feed elif mode == "display": model.load_weights('model.h5') # prevents openCL usage and unnecessary logging messages cv2.ocl.setUseOpenCL(False) # dictionary which assigns each label an emotion (alphabetical order) emotion_dict = {0: "Angry", 1: "Disgusted", 2: "Fearful", 3: "Happy", 4: "Neutral", 5: "Sad", 6: "Surprised"} # start the webcam feed cap = cv2.VideoCapture(1) # 平滑配置:可根据需求调整阈值,数值越大越稳定,但是响应延迟越高 SMOOTH_THRESHOLD = 5 last_emotion = None consecutive_count = 0 stable_emotion = None while True: # Find haar cascade to draw bounding box around face ret, frame = cap.read() if not ret: break facecasc = cv2.CascadeClassifier('haarcascade_frontalface_default.xml') gray = cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) faces = facecasc.detectMultiScale(gray,scaleFactor=1.3, minNeighbors=5) for (x, y, w, h) in faces: cv2.rectangle(frame, (x, y-50), (x+w, y+h+10), (255, 0, 0), 2) roi_gray = gray[y:y + h, x:x + w] cropped_img = np.expand_dims(np.expand_dims(cv2.resize(roi_gray, (48, 48)), -1), 0) prediction = model.predict(cropped_img) maxindex = int(np.argmax(prediction)) text = emotion_dict[maxindex] if ("Sad" in text) or ("Angry" in text) or ("Disgusted" in text): text = "Sad" # 新增平滑逻辑 if text == last_emotion: consecutive_count += 1 else: consecutive_count = 1 last_emotion = text # 连续命中阈值才更新稳定展示的情绪 if consecutive_count >= SMOOTH_THRESHOLD: stable_emotion = text # 仅展示稳定的情绪结果 if stable_emotion is not None and ("Happy" in stable_emotion or "Sad" in stable_emotion): cv2.putText(frame, stable_emotion, (x+20, y-60), cv2.FONT_HERSHEY_SIMPLEX, 1, (255, 255, 255), 2, cv2.LINE_AA) cv2.imshow('Video', cv2.resize(frame,(1600,960),interpolation = cv2.INTER_CUBIC)) if cv2.waitKey(1) & 0xFF == ord('q'): break cap.release() cv2.destroyAllWindows()
补充说明
上述代码默认适配单人脸识别场景,如果需要支持多人脸识别,可将三个状态变量改为字典结构,以人脸的坐标特征作为key,单独维护每个检测到的人脸的连续计数即可。
内容的提问来源于stack exchange,提问作者angel_30
相关产品推荐
相关产品推荐

