Flask框架人脸情绪识别Web应用音视频文同步问题咨询
卡顿根因
- 阻塞逻辑位置错误:
playsound为同步阻塞方法,放在视频帧生成主循环中时,播放音频会直接阻断帧读取、推理、推流全流程,直接造成视频卡顿。 - 事件触发频率过高:消费级摄像头默认输出25-30帧/秒,逐帧推理+触发音频会导致1秒内重复触发数十次播放请求,资源占满后自然出现音画文本不同步问题。
- 架构逻辑不合理:音频播放属于前端交互层逻辑,不应放在后端推流线程中执行。后端仅需返回识别结果,音画同步调度由前端实现即可。
3秒采样+音文同步实现方案
核心逻辑:后端视频流保持正常逐帧推送保证流畅度,单独缓存识别结果,每3秒完成一次采样推理更新结果,对外提供结果查询接口;前端定时拉取结果,同步更新文本、触发音频播放,全程不阻塞视频流传输。
后端app.py修改
移除原有playsound相关导入与阻塞逻辑,增加全局变量缓存识别结果、控制3秒采样间隔,新增结果查询接口,开启多线程模式避免请求阻塞。
from flask import Flask, render_template, Response, jsonify import cv2 import numpy as np import time from tensorflow.keras.models import model_from_json from tensorflow.keras.preprocessing import image # 加载模型 model = model_from_json(open(r'C:\Users\HP\emotion_model.json', 'r').read()) model.load_weights(r'C:\Users\HP\emotion_model.h5') face_haar_cascade = cv2.CascadeClassifier(r'C:\Users\HP\haarcascade_frontalface_default.xml') app = Flask(__name__) camera = cv2.VideoCapture(0) # 全局缓存配置 latest_emotion = "Neutral" last_predict_time = 0 PREDICT_INTERVAL = 3 # 采样间隔3秒 def gen_frames(): global latest_emotion, last_predict_time while True: success, frame = camera.read() if not success: break gray_img= cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY) faces_detected = face_haar_cascade.detectMultiScale(gray_img, 1.32, 5) current_time = time.time() for (x,y,w,h) in faces_detected: cv2.rectangle(frame,(x,y),(x+w,y+h),(255,0,0),thickness=4) # 仅到采样间隔才跑推理,其余帧仅绘制已缓存的结果 if current_time - last_predict_time >= PREDICT_INTERVAL: roi_gray=gray_img[y:y+w,x:x+h] roi_gray=cv2.resize(roi_gray,(48,48)) img_pixels = image.img_to_array(roi_gray) img_pixels = np.expand_dims(img_pixels, axis = 0) img_pixels /= 255 predictions = model.predict(img_pixels, verbose=0) max_index = np.argmax(predictions[0]) emotions = ['Angry', 'Disgust', 'Fear', 'Happy', 'Sad', 'Surprise', 'Neutral'] latest_emotion = emotions[max_index] last_predict_time = current_time # 所有帧均绘制表情标签,保证画面显示连续 cv2.putText(frame, latest_emotion, (int(x), int(y)), cv2.FONT_HERSHEY_SIMPLEX, 1, (0,0,255), 2) ret, buffer = cv2.imencode('.jpg', frame) frame = buffer.tobytes() yield (b'--frame\r\n' b'Content-Type: image/jpeg\r\n\r\n' + frame + b'\r\n') # 前端轮询用的结果接口 @app.route('/get_latest_emotion') def get_latest_emotion(): return jsonify({"emotion": latest_emotion}) @app.route('/video_feed') def video_feed(): return Response(gen_frames(), mimetype='multipart/x-mixed-replace; boundary=frame') @app.route('/') def index(): return render_template('index.html') if __name__ == '__main__': app.run(debug=True, threaded=True)
前端index.html修改
提前在项目static目录下创建audio文件夹,将所有表情对应mp3文件放入(文件名与情绪名小写保持一致,如happy.mp3、angry.mp3),新增文本显示区、音频播放器,加定时轮询逻辑,拿到结果后同步更新文本、播放音频,增加重复播放拦截。
<!doctype html> <html lang="en"> <head> <meta charset="utf-8"> <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no"> <meta http-equiv="X-UA-Compatible" content="IE=edge,chrome=1"> <link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css" integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous"> <title>Real Time Emotion Detection</title> <style> #emotion-text { color: white; font-size: 2rem; font-family: verdana; margin: 1rem 0; } </style> </head> <body style="background-color:#002147;"> <div class="parallax-content baner-content" id="home"> <div class="container"> <div class="row"> <div class="col-lg-8 offset-lg-2"> <h3 class="mt-5"><font color="white" style="font-family:verdana;font-size:300%;"><center>Real-Time Emotion Detection</center></font></h3> <center><img src="{{ url_for('video_feed') }}" width="80%"></center> <center><div id="emotion-text">检测中...</div></center> </div> </div> </div> </div> <audio id="audio-player" preload="auto"></audio> <script> const emotionText = document.getElementById('emotion-text'); const audioPlayer = document.getElementById('audio-player'); let lastPlayedEmotion = ''; // 每3秒拉取一次最新结果 setInterval(async () => { const res = await fetch('/get_latest_emotion'); const data = await res.json(); const emotion = data.emotion; // 同步更新文本 emotionText.textContent = `当前表情:${emotion}`; // 情绪变化时才播放音频,避免重复触发 if (emotion !== lastPlayedEmotion) { audioPlayer.src = `/static/audio/${emotion.toLowerCase()}.mp3`; audioPlayer.play().catch(()=>{}); lastPlayedEmotion = emotion; } }, 3000); </script> </body> </html>
优化说明
- 视频流传输全程不被推理、音频逻辑阻塞,帧率稳定无卡顿。
- 3秒一次的推理频率大幅降低算力占用,低配设备也能流畅运行。
- 前端拿到结果后同时触发文本更新和音频播放,天然同步无延迟。
- 相同情绪不重复触发播放,避免音频叠加错乱。
浏览器默认禁止未交互页面自动播放音频,首次打开页面后点击任意位置即可正常触发音频播放。
内容的提问来源于stack exchange,提问作者Amith Kotian
相关产品推荐
相关产品推荐

