基于Streamlit实现按住空格键录音转文字的功能开发求助
问题
我正在开发一款简单的Streamlit网页应用,需要实现以下功能:
- 用户按住空格键并说话
- 用户松开空格键后,将所说内容转换为文字并返回至UI界面
目前已实现点击按钮转文字的功能,但语音识别有时会错过开头或过早截断。尝试用keyboard库监听空格键时,应用直接挂起,无法正常运行。现有按钮版代码如下:
import streamlit as st import speech_recognition as sr #(SpeechRecognition in pypi) def transcribe_speech(): r = sr.Recognizer() with sr.Microphone() as source: r.adjust_for_ambient_noise(source) with st.spinner("Listening..."): audio = r.listen(source) st.write("Transcribing...") try: text = r.recognize_google(audio) return text except sr.UnknownValueError: st.write("Could not understand audio") except sr.RequestError as e: st.write("Could not request results from Google Speech Recognition service; {0}".format(e)) ready_button = st.button("TALK TO ME", key='ready_button') if ready_button: text = transcribe_speech() if text: st.write(f"You said: {text}")
尝试的空格键监听代码(导致应用挂起):
import keyboard import time def print_current_time(): current_time = time.strftime("%H:%M:%S") st.write(f"Current time: {current_time}") space_pressed = False while True: if keyboard.is_pressed(" "): st.write("if keyboard.is_pressed") if not space_pressed: st.write("if not space_pressed") space_pressed = True print_current_time() else: space_pressed = False time.sleep(1)
需要实现按住空格键录音、松开后转文字的功能。
解决方案
核心思路
Streamlit是Web应用,不能用本地的keyboard库监听按键——需要通过前端JavaScript监听浏览器的键盘事件,再将状态传递给后端。同时,录音操作是阻塞的,需要用线程来避免卡住应用界面。
完整实现代码
import streamlit as st import speech_recognition as sr import pyaudio import wave import threading import time # 初始化会话状态 if "recording" not in st.session_state: st.session_state.recording = False if "audio_file" not in st.session_state: st.session_state.audio_file = "temp_recording.wav" if "transcript" not in st.session_state: st.session_state.transcript = "" # 录音参数设置 FORMAT = pyaudio.paInt16 CHANNELS = 1 RATE = 44100 CHUNK = 1024 def start_recording(): """开始录音,保存到临时文件""" p = pyaudio.PyAudio() stream = p.open(format=FORMAT, channels=CHANNELS, rate=RATE, input=True, frames_per_buffer=CHUNK) frames = [] while st.session_state.recording: data = stream.read(CHUNK) frames.append(data) # 停止录音 stream.stop_stream() stream.close() p.terminate() # 保存音频文件 wf = wave.open(st.session_state.audio_file, 'wb') wf.setnchannels(CHANNELS) wf.setsampwidth(p.get_sample_size(FORMAT)) wf.setframerate(RATE) wf.writeframes(b''.join(frames)) wf.close() def transcribe_audio(): """将录音文件转文字""" r = sr.Recognizer() with sr.AudioFile(st.session_state.audio_file) as source: audio = r.record(source) try: text = r.recognize_google(audio) st.session_state.transcript = text except sr.UnknownValueError: st.session_state.transcript = "无法识别音频内容" except sr.RequestError as e: st.session_state.transcript = f"语音识别服务请求失败: {e}" # 注入JavaScript监听空格键事件 st.markdown(""" <script> document.addEventListener('keydown', function(e) { if (e.code === 'Space' && !e.repeat) { // 按下空格键,发送开始录音信号 Streamlit.setComponentValue("start_recording", true); } }); document.addEventListener('keyup', function(e) { if (e.code === 'Space') { // 松开空格键,发送停止录音信号 Streamlit.setComponentValue("stop_recording", true); } }); </script> """, unsafe_allow_html=True) # 处理前端传来的信号 if st.experimental_get_query_params().get("start_recording") == ["true"]: if not st.session_state.recording: st.session_state.recording = True # 启动录音线程 threading.Thread(target=start_recording).start() st.experimental_set_query_params() # 清空参数 st.rerun() if st.experimental_get_query_params().get("stop_recording") == ["true"]: if st.session_state.recording: st.session_state.recording = False time.sleep(0.5) # 等待录音线程完成文件保存 transcribe_audio() st.experimental_set_query_params() # 清空参数 st.rerun() # UI展示 st.title("按住空格键录音") if st.session_state.recording: st.warning("正在录音... 松开空格键结束") else: st.info("按住空格键开始说话,松开后自动转文字") if st.session_state.transcript: st.subheader("识别结果:") st.write(st.session_state.transcript)
关键说明
- 前端按键监听:通过注入JavaScript代码,监听浏览器的
keydown和keyup事件,当空格键按下/松开时,通过Streamlit.setComponentValue传递状态给后端。 - 线程化录音:录音操作是阻塞的,用
threading.Thread启动录音线程,避免卡住Streamlit的UI渲染。 - 会话状态管理:用
st.session_state跟踪录音状态、临时音频文件路径和识别结果,确保页面刷新时状态不丢失。 - 音频转文字:录音停止后,读取临时WAV文件,用
speech_recognition调用Google语音识别接口转换文字。
依赖安装
需要安装额外依赖:
pip install streamlit speechrecognition pyaudio
注意:pyaudio在部分系统可能需要手动安装依赖(如Ubuntu需先安装portaudio19-dev)。
内容的提问来源于stack exchange,提问作者GivenX
相关产品推荐
相关产品推荐

