You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Flask框架人脸情绪识别Web应用音视频文同步问题咨询

卡顿根因
  • 阻塞逻辑位置错误:playsound为同步阻塞方法,放在视频帧生成主循环中时,播放音频会直接阻断帧读取、推理、推流全流程,直接造成视频卡顿。
  • 事件触发频率过高:消费级摄像头默认输出25-30帧/秒,逐帧推理+触发音频会导致1秒内重复触发数十次播放请求,资源占满后自然出现音画文本不同步问题。
  • 架构逻辑不合理:音频播放属于前端交互层逻辑,不应放在后端推流线程中执行。后端仅需返回识别结果,音画同步调度由前端实现即可。
3秒采样+音文同步实现方案

核心逻辑:后端视频流保持正常逐帧推送保证流畅度,单独缓存识别结果,每3秒完成一次采样推理更新结果,对外提供结果查询接口;前端定时拉取结果,同步更新文本、触发音频播放,全程不阻塞视频流传输。

后端app.py修改

移除原有playsound相关导入与阻塞逻辑,增加全局变量缓存识别结果、控制3秒采样间隔,新增结果查询接口,开启多线程模式避免请求阻塞。

from flask import Flask, render_template, Response, jsonify
import cv2
import numpy as np
import time
from tensorflow.keras.models import model_from_json  
from tensorflow.keras.preprocessing import image  
  

# 加载模型
model = model_from_json(open(r'C:\Users\HP\emotion_model.json', 'r').read())  
model.load_weights(r'C:\Users\HP\emotion_model.h5')
face_haar_cascade = cv2.CascadeClassifier(r'C:\Users\HP\haarcascade_frontalface_default.xml')

app = Flask(__name__)
camera = cv2.VideoCapture(0)

# 全局缓存配置
latest_emotion = "Neutral"
last_predict_time = 0
PREDICT_INTERVAL = 3  # 采样间隔3秒

def gen_frames():
    global latest_emotion, last_predict_time
    while True:
        success, frame = camera.read()
        if not success:
            break
        gray_img= cv2.cvtColor(frame, cv2.COLOR_BGR2GRAY)  
        faces_detected = face_haar_cascade.detectMultiScale(gray_img, 1.32, 5)  
        current_time = time.time()
        
        for (x,y,w,h) in faces_detected:
            cv2.rectangle(frame,(x,y),(x+w,y+h),(255,0,0),thickness=4)
            # 仅到采样间隔才跑推理,其余帧仅绘制已缓存的结果
            if current_time - last_predict_time >= PREDICT_INTERVAL:
                roi_gray=gray_img[y:y+w,x:x+h]
                roi_gray=cv2.resize(roi_gray,(48,48))  
                img_pixels = image.img_to_array(roi_gray)  
                img_pixels = np.expand_dims(img_pixels, axis = 0)  
                img_pixels /= 255  
                predictions = model.predict(img_pixels, verbose=0)
                max_index = np.argmax(predictions[0])  
                emotions = ['Angry', 'Disgust', 'Fear', 'Happy', 'Sad', 'Surprise', 'Neutral'] 
                latest_emotion = emotions[max_index]
                last_predict_time = current_time
            # 所有帧均绘制表情标签,保证画面显示连续
            cv2.putText(frame, latest_emotion, (int(x), int(y)), cv2.FONT_HERSHEY_SIMPLEX, 1, (0,0,255), 2)

        ret, buffer = cv2.imencode('.jpg', frame)
        frame = buffer.tobytes()
        yield (b'--frame\r\n'
               b'Content-Type: image/jpeg\r\n\r\n' + frame + b'\r\n')

# 前端轮询用的结果接口
@app.route('/get_latest_emotion')
def get_latest_emotion():
    return jsonify({"emotion": latest_emotion})

@app.route('/video_feed')
def video_feed():
    return Response(gen_frames(), mimetype='multipart/x-mixed-replace; boundary=frame')

@app.route('/')
def index():
    return render_template('index.html')

if __name__ == '__main__':
    app.run(debug=True, threaded=True)

前端index.html修改

提前在项目static目录下创建audio文件夹,将所有表情对应mp3文件放入(文件名与情绪名小写保持一致,如happy.mp3、angry.mp3),新增文本显示区、音频播放器,加定时轮询逻辑,拿到结果后同步更新文本、播放音频,增加重复播放拦截。

<!doctype html>
<html lang="en">
<head>
    <meta charset="utf-8">
    <meta name="viewport" content="width=device-width, initial-scale=1, shrink-to-fit=no">
    <meta http-equiv="X-UA-Compatible" content="IE=edge,chrome=1">
    <link rel="stylesheet" href="https://stackpath.bootstrapcdn.com/bootstrap/4.1.3/css/bootstrap.min.css"
          integrity="sha384-MCw98/SFnGE8fJT3GXwEOngsV7Zt27NXFoaoApmYm81iuXoPkFOJwJ8ERdknLPMO" crossorigin="anonymous">
    <title>Real Time Emotion Detection</title>
    <style>
        #emotion-text {
            color: white;
            font-size: 2rem;
            font-family: verdana;
            margin: 1rem 0;
        }
    </style>
</head>
<body style="background-color:#002147;">
    <div class="parallax-content baner-content" id="home">
        <div class="container">
            <div class="row">
                <div class="col-lg-8  offset-lg-2">
                    <h3 class="mt-5"><font color="white" style="font-family:verdana;font-size:300%;"><center>Real-Time Emotion Detection</center></font></h3>
                    <center><img src="{{ url_for('video_feed') }}" width="80%"></center>
                    <center><div id="emotion-text">检测中...</div></center>
                </div>
            </div>
        </div>
    </div>
    <audio id="audio-player" preload="auto"></audio>
    <script>
        const emotionText = document.getElementById('emotion-text');
        const audioPlayer = document.getElementById('audio-player');
        let lastPlayedEmotion = '';
        // 每3秒拉取一次最新结果
        setInterval(async () => {
            const res = await fetch('/get_latest_emotion');
            const data = await res.json();
            const emotion = data.emotion;
            // 同步更新文本
            emotionText.textContent = `当前表情:${emotion}`;
            // 情绪变化时才播放音频,避免重复触发
            if (emotion !== lastPlayedEmotion) {
                audioPlayer.src = `/static/audio/${emotion.toLowerCase()}.mp3`;
                audioPlayer.play().catch(()=>{});
                lastPlayedEmotion = emotion;
            }
        }, 3000);
    </script>
</body>
</html>
优化说明
  • 视频流传输全程不被推理、音频逻辑阻塞,帧率稳定无卡顿。
  • 3秒一次的推理频率大幅降低算力占用,低配设备也能流畅运行。
  • 前端拿到结果后同时触发文本更新和音频播放,天然同步无延迟。
  • 相同情绪不重复触发播放,避免音频叠加错乱。

浏览器默认禁止未交互页面自动播放音频,首次打开页面后点击任意位置即可正常触发音频播放。

内容的提问来源于stack exchange,提问作者Amith Kotian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.29 01:18:21