如何在dash-devices中使用客户端回调实现麦克风语音转录
实现方案
你原来的代码核心问题是直接在服务端调用sr.Microphone(),只能读取服务端本地的麦克风设备,部署到远程服务器后完全无法采集用户端的语音。正确的实现路径是通过浏览器端JS调用MediaRecorder API采集用户本地音频,再将音频数据回传给Python后端处理转录,结合dash-devices自带的多用户状态同步特性,就能实现两名用户同时查看转录结果的需求。
完整改造代码
from dash_devices.dependencies import Input, Output, State import dash_html_components as html import dash_core_components as dcc import speech_recognition as sr import base64 import io app = dash_devices.Dash(__name__) app.config.suppress_callback_exceptions = True r = sr.Recognizer() app.layout = html.Div([ html.Div("Transcription", id='transcription'), html.Button(id='listen-pause', children='Record Message'), # 新增store存储录音状态和音频数据 dcc.Store(id='is-recording', data=False), dcc.Store(id='audio-data', data='') ]) # 客户端回调:JS实现浏览器端录音逻辑 app.clientside_callback( """ function(n_clicks, is_recording) { if (n_clicks === 0) return [false, '', 'Record Message']; const mediaRecorder = window.mediaRecorder || null; // 切换录音状态 if (!is_recording) { // 请求麦克风权限并开始录音 navigator.mediaDevices.getUserMedia({ audio: true }) .then(stream => { const recorder = new MediaRecorder(stream, { mimeType: 'audio/webm' }); window.mediaRecorder = recorder; const audioChunks = []; recorder.ondataavailable = event => audioChunks.push(event.data); recorder.onstop = () => { const audioBlob = new Blob(audioChunks, { type: 'audio/webm' }); const reader = new FileReader(); reader.readAsDataURL(audioBlob); reader.onloadend = () => { // 音频转base64存入store const base64Audio = reader.result.split(',')[1]; document.getElementById('audio-data').data = base64Audio; }; stream.getTracks().forEach(track => track.stop()); }; recorder.start(); // 3秒后自动停止,和原代码逻辑一致 setTimeout(() => { if(recorder.state !== 'inactive') recorder.stop() }, 3000); }) .catch(err => { console.error('麦克风调用失败', err); document.getElementById('transcription').innerText = '请授予麦克风访问权限'; }); return [true, '', 'Recording...']; } else { // 手动停止录音 if (mediaRecorder && mediaRecorder.state !== 'inactive') { mediaRecorder.stop(); } return [false, '', 'Record Message']; } } """, [Output('is-recording', 'data'), Output('audio-data', 'data'), Output('listen-pause', 'children')], [Input('listen-pause', 'n_clicks')], [State('is-recording', 'data')], prevent_initial_call=True ) # 服务端回调:处理音频转写,结果自动同步给所有在线用户 @app.callback( Output(component_id='transcription', component_property='children'), Input(component_id='audio-data', component_property='data'), prevent_initial_call=True ) def transcribe_speech(base64_audio): if not base64_audio: return "" try: # 解码base64音频为字节流 audio_bytes = base64.b64decode(base64_audio) audio_file = io.BytesIO(audio_bytes) with sr.AudioFile(audio_file) as source: audio_data = r.record(source) # 可修改language参数适配其他语言,比如英文填入'en-US' transcript = r.recognize_google(audio_data, language='zh-CN') return transcript except sr.UnknownValueError: return "Could not parse input" except Exception as e: return f"处理出错:{str(e)}" if __name__ == '__main__': app.run_server(debug=True, host='0.0.0.0', port=5000)
注意事项
- 浏览器安全限制要求必须在HTTPS环境或者localhost下才能调用麦克风权限,远程部署时需要配置SSL证书
- 可自行修改JS代码中的
setTimeout参数调整最长录音时长,也可以去掉自动停止逻辑,完全由用户点击按钮控制录音启停 - 如需兼容更多浏览器,可调整MediaRecorder的mimeType参数,不同浏览器支持的音频格式略有差异
内容的提问来源于stack exchange,提问作者Nick
相关产品推荐
相关产品推荐

