You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Flask+TensorFlow实现浏览器实时语音识别功能需求

实现Flask+TensorFlow实时语音识别并返回前端

Flask后端修改(集成TensorFlow语音识别)

首先安装依赖:

pip install flask tensorflow soundfile numpy

修改后的后端代码,负责接收前端音频流、预处理、模型识别并返回结果:

from flask import Flask, request, jsonify, render_template
import tensorflow as tf
import soundfile as sf
import numpy as np
import io

app = Flask(__name__)

# 加载你的TensorFlow语音识别模型
# 替换为实际模型路径,若使用预训练模型(如Whisper TF版),需调整加载逻辑
model = tf.keras.models.load_model('path/to/your/asr_model')

# 音频预处理:转换为单声道、16000采样率,适配模型输入要求
def preprocess_audio(audio_bytes):
    # 从字节流读取音频
    with io.BytesIO(audio_bytes) as f:
        data, sample_rate = sf.read(f)
    
    # 转单声道
    if len(data.shape) > 1:
        data = np.mean(data, axis=1)
    
    # 重采样到16000Hz
    if sample_rate != 16000:
        data = tf.audio.resample(data, sample_rate, 16000).numpy()
    
    # 统一音频长度(示例为1秒,根据你的模型输入调整)
    target_len = 16000
    if len(data) < target_len:
        data = np.pad(data, (0, target_len - len(data)), mode='constant')
    else:
        data = data[:target_len]
    
    # 归一化(匹配模型训练时的预处理逻辑)
    data = data / np.max(np.abs(data))
    return data.reshape(1, -1)

# 语音识别逻辑,根据模型输出调整解码方式
def recognize_text(audio_data):
    # 模型预测
    logits = model.predict(audio_data, verbose=0)
    
    # 示例:CTC贪婪解码(如果模型用CTC损失训练)
    # 替换为你的标签映射表,比如包含所有汉字/字母的映射
    label_map = {0: '', 1: '你', 2: '好', 3: '世', 4: '界', 5: '我', 6: '是'}
    predicted_ids = tf.argmax(logits, axis=-1).numpy()[0]
    
    # 拼接成文本
    text = ''.join([label_map[id] for id in predicted_ids if id != 0])
    return text.strip()

@app.route('/')
def home():
    return render_template('index.html')

@app.route('/api/recognize', methods=['POST'])
def handle_recognition():
    if 'audio' not in request.files:
        return jsonify({'error': '未上传音频文件'}), 400
    
    audio_file = request.files['audio']
    try:
        audio_bytes = audio_file.read()
        processed_audio = preprocess_audio(audio_bytes)
        result_text = recognize_text(processed_audio)
        return jsonify({'result': result_text})
    except Exception as e:
        return jsonify({'error': str(e)}), 500

if __name__ == '__main__':
    app.run(debug=True, host='0.0.0.0')

前端页面适配(生成符合要求的音频流)

将前端HTML文件放在Flask的templates目录下,确保录制的音频是单声道、16000采样率的WAV格式,并实时发送到后端:

<!DOCTYPE html>
<html lang="zh-CN">
<head>
    <meta charset="UTF-8">
    <title>实时语音识别</title>
</head>
<body>
    <button id="startRecord">开始录制</button>
    <button id="stopRecord" disabled>停止录制</button>
    <div style="margin-top: 20px;">识别结果:<span id="resultText"></span></div>

    <script>
        let mediaRecorder;
        const startBtn = document.getElementById('startRecord');
        const stopBtn = document.getElementById('stopRecord');
        const resultSpan = document.getElementById('resultText');

        // 音频录制配置:严格匹配要求
        const audioConstraints = {
            audio: {
                channelCount: 1,
                sampleRate: 16000,
                sampleSize: 16,
                echoCancellation: true,
                noiseSuppression: true
            }
        };

        startBtn.addEventListener('click', async () => {
            const stream = await navigator.mediaDevices.getUserMedia(audioConstraints);
            // 指定录制格式为WAV
            mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/wav' });

            // 每500ms发送一次音频片段(可根据实时性需求调整)
            mediaRecorder.start(500);

            mediaRecorder.ondataavailable = async (e) => {
                if (e.data.size > 0) {
                    await sendAudioToBackend(e.data);
                }
            };

            startBtn.disabled = true;
            stopBtn.disabled = false;
        });

        stopBtn.addEventListener('click', () => {
            mediaRecorder.stop();
            startBtn.disabled = false;
            stopBtn.disabled = true;
        });

        // 发送音频到后端
        async function sendAudioToBackend(audioBlob) {
            const formData = new FormData();
            formData.append('audio', audioBlob, 'temp.wav');

            try {
                const res = await fetch('/api/recognize', {
                    method: 'POST',
                    body: formData
                });
                const data = await res.json();
                if (data.result) {
                    resultSpan.textContent = data.result;
                } else if (data.error) {
                    resultSpan.textContent = `错误:${data.error}`;
                }
            } catch (err) {
                resultSpan.textContent = `请求失败:${err.message}`;
            }
        }
    </script>
</body>
</html>

关键注意事项

  • 模型适配:如果需要识别连续自然语言,建议使用TensorFlow版的Whisper模型(可通过TensorFlow Hub加载),上述示例的CTC解码仅适用于特定训练的模型,需根据你的模型输出调整解码逻辑。
  • 音频一致性:必须保证前端录制和后端预处理后的音频参数完全匹配(单声道、16000采样率),否则模型无法正常工作。
  • 实时性优化:若追求更低延迟,可将HTTP POST替换为WebSocket协议,减少请求开销;同时调整前端录制片段的时长,平衡实时性和识别准确率。
  • 依赖问题:Linux系统下安装soundfile可能需要先安装系统依赖:sudo apt-get install libsndfile1-dev。

内容的提问来源于stack exchange,提问作者mahendra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 17:33:25