基于Flask+TensorFlow实现浏览器实时语音识别功能需求
实现Flask+TensorFlow实时语音识别并返回前端
Flask后端修改(集成TensorFlow语音识别)
首先安装依赖:
pip install flask tensorflow soundfile numpy
修改后的后端代码,负责接收前端音频流、预处理、模型识别并返回结果:
from flask import Flask, request, jsonify, render_template import tensorflow as tf import soundfile as sf import numpy as np import io app = Flask(__name__) # 加载你的TensorFlow语音识别模型 # 替换为实际模型路径,若使用预训练模型(如Whisper TF版),需调整加载逻辑 model = tf.keras.models.load_model('path/to/your/asr_model') # 音频预处理:转换为单声道、16000采样率,适配模型输入要求 def preprocess_audio(audio_bytes): # 从字节流读取音频 with io.BytesIO(audio_bytes) as f: data, sample_rate = sf.read(f) # 转单声道 if len(data.shape) > 1: data = np.mean(data, axis=1) # 重采样到16000Hz if sample_rate != 16000: data = tf.audio.resample(data, sample_rate, 16000).numpy() # 统一音频长度(示例为1秒,根据你的模型输入调整) target_len = 16000 if len(data) < target_len: data = np.pad(data, (0, target_len - len(data)), mode='constant') else: data = data[:target_len] # 归一化(匹配模型训练时的预处理逻辑) data = data / np.max(np.abs(data)) return data.reshape(1, -1) # 语音识别逻辑,根据模型输出调整解码方式 def recognize_text(audio_data): # 模型预测 logits = model.predict(audio_data, verbose=0) # 示例:CTC贪婪解码(如果模型用CTC损失训练) # 替换为你的标签映射表,比如包含所有汉字/字母的映射 label_map = {0: '', 1: '你', 2: '好', 3: '世', 4: '界', 5: '我', 6: '是'} predicted_ids = tf.argmax(logits, axis=-1).numpy()[0] # 拼接成文本 text = ''.join([label_map[id] for id in predicted_ids if id != 0]) return text.strip() @app.route('/') def home(): return render_template('index.html') @app.route('/api/recognize', methods=['POST']) def handle_recognition(): if 'audio' not in request.files: return jsonify({'error': '未上传音频文件'}), 400 audio_file = request.files['audio'] try: audio_bytes = audio_file.read() processed_audio = preprocess_audio(audio_bytes) result_text = recognize_text(processed_audio) return jsonify({'result': result_text}) except Exception as e: return jsonify({'error': str(e)}), 500 if __name__ == '__main__': app.run(debug=True, host='0.0.0.0')
前端页面适配(生成符合要求的音频流)
将前端HTML文件放在Flask的templates目录下,确保录制的音频是单声道、16000采样率的WAV格式,并实时发送到后端:
<!DOCTYPE html> <html lang="zh-CN"> <head> <meta charset="UTF-8"> <title>实时语音识别</title> </head> <body> <button id="startRecord">开始录制</button> <button id="stopRecord" disabled>停止录制</button> <div style="margin-top: 20px;">识别结果:<span id="resultText"></span></div> <script> let mediaRecorder; const startBtn = document.getElementById('startRecord'); const stopBtn = document.getElementById('stopRecord'); const resultSpan = document.getElementById('resultText'); // 音频录制配置:严格匹配要求 const audioConstraints = { audio: { channelCount: 1, sampleRate: 16000, sampleSize: 16, echoCancellation: true, noiseSuppression: true } }; startBtn.addEventListener('click', async () => { const stream = await navigator.mediaDevices.getUserMedia(audioConstraints); // 指定录制格式为WAV mediaRecorder = new MediaRecorder(stream, { mimeType: 'audio/wav' }); // 每500ms发送一次音频片段(可根据实时性需求调整) mediaRecorder.start(500); mediaRecorder.ondataavailable = async (e) => { if (e.data.size > 0) { await sendAudioToBackend(e.data); } }; startBtn.disabled = true; stopBtn.disabled = false; }); stopBtn.addEventListener('click', () => { mediaRecorder.stop(); startBtn.disabled = false; stopBtn.disabled = true; }); // 发送音频到后端 async function sendAudioToBackend(audioBlob) { const formData = new FormData(); formData.append('audio', audioBlob, 'temp.wav'); try { const res = await fetch('/api/recognize', { method: 'POST', body: formData }); const data = await res.json(); if (data.result) { resultSpan.textContent = data.result; } else if (data.error) { resultSpan.textContent = `错误:${data.error}`; } } catch (err) { resultSpan.textContent = `请求失败:${err.message}`; } } </script> </body> </html>
关键注意事项
- 模型适配:如果需要识别连续自然语言,建议使用TensorFlow版的Whisper模型(可通过TensorFlow Hub加载),上述示例的CTC解码仅适用于特定训练的模型,需根据你的模型输出调整解码逻辑。
- 音频一致性:必须保证前端录制和后端预处理后的音频参数完全匹配(单声道、16000采样率),否则模型无法正常工作。
- 实时性优化:若追求更低延迟,可将HTTP POST替换为WebSocket协议,减少请求开销;同时调整前端录制片段的时长,平衡实时性和识别准确率。
- 依赖问题:Linux系统下安装
soundfile可能需要先安装系统依赖:sudo apt-get install libsndfile1-dev。
内容的提问来源于stack exchange,提问作者mahendra
相关产品推荐
相关产品推荐

