You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Streamlit实现按住空格键录音转文字的功能开发求助

问题

我正在开发一款简单的Streamlit网页应用,需要实现以下功能:

  • 用户按住空格键并说话
  • 用户松开空格键后,将所说内容转换为文字并返回至UI界面

目前已实现点击按钮转文字的功能,但语音识别有时会错过开头或过早截断。尝试用keyboard库监听空格键时,应用直接挂起,无法正常运行。现有按钮版代码如下:

import streamlit as st
import speech_recognition as sr #(SpeechRecognition in pypi) 

def transcribe_speech():
    r = sr.Recognizer()
    with sr.Microphone() as source:
        r.adjust_for_ambient_noise(source)
        with st.spinner("Listening..."):
            audio = r.listen(source)
            st.write("Transcribing...")
            try:
                text = r.recognize_google(audio)
                return text
            except sr.UnknownValueError:
                st.write("Could not understand audio")
            except sr.RequestError as e:
                st.write("Could not request results from Google Speech Recognition service; {0}".format(e))

ready_button = st.button("TALK TO ME", key='ready_button')

if ready_button:
    text = transcribe_speech()
    if text:
        st.write(f"You said: {text}")

尝试的空格键监听代码(导致应用挂起):

import keyboard
import time

def print_current_time():
    current_time = time.strftime("%H:%M:%S")
    st.write(f"Current time: {current_time}")

space_pressed = False

while True:
    if keyboard.is_pressed(" "):
        st.write("if keyboard.is_pressed")
        if not space_pressed:
            st.write("if not space_pressed")
            space_pressed = True
            print_current_time()
    else:
        space_pressed = False
    time.sleep(1)

需要实现按住空格键录音、松开后转文字的功能。

解决方案

核心思路

Streamlit是Web应用,不能用本地的keyboard库监听按键——需要通过前端JavaScript监听浏览器的键盘事件,再将状态传递给后端。同时,录音操作是阻塞的,需要用线程来避免卡住应用界面。

完整实现代码

import streamlit as st
import speech_recognition as sr
import pyaudio
import wave
import threading
import time

# 初始化会话状态
if "recording" not in st.session_state:
    st.session_state.recording = False
if "audio_file" not in st.session_state:
    st.session_state.audio_file = "temp_recording.wav"
if "transcript" not in st.session_state:
    st.session_state.transcript = ""

# 录音参数设置
FORMAT = pyaudio.paInt16
CHANNELS = 1
RATE = 44100
CHUNK = 1024

def start_recording():
    """开始录音,保存到临时文件"""
    p = pyaudio.PyAudio()
    stream = p.open(format=FORMAT, channels=CHANNELS,
                    rate=RATE, input=True,
                    frames_per_buffer=CHUNK)
    frames = []
    
    while st.session_state.recording:
        data = stream.read(CHUNK)
        frames.append(data)
    
    # 停止录音
    stream.stop_stream()
    stream.close()
    p.terminate()
    
    # 保存音频文件
    wf = wave.open(st.session_state.audio_file, 'wb')
    wf.setnchannels(CHANNELS)
    wf.setsampwidth(p.get_sample_size(FORMAT))
    wf.setframerate(RATE)
    wf.writeframes(b''.join(frames))
    wf.close()

def transcribe_audio():
    """将录音文件转文字"""
    r = sr.Recognizer()
    with sr.AudioFile(st.session_state.audio_file) as source:
        audio = r.record(source)
        try:
            text = r.recognize_google(audio)
            st.session_state.transcript = text
        except sr.UnknownValueError:
            st.session_state.transcript = "无法识别音频内容"
        except sr.RequestError as e:
            st.session_state.transcript = f"语音识别服务请求失败: {e}"

# 注入JavaScript监听空格键事件
st.markdown("""
<script>
document.addEventListener('keydown', function(e) {
    if (e.code === 'Space' && !e.repeat) {
        // 按下空格键,发送开始录音信号
        Streamlit.setComponentValue("start_recording", true);
    }
});

document.addEventListener('keyup', function(e) {
    if (e.code === 'Space') {
        // 松开空格键,发送停止录音信号
        Streamlit.setComponentValue("stop_recording", true);
    }
});
</script>
""", unsafe_allow_html=True)

# 处理前端传来的信号
if st.experimental_get_query_params().get("start_recording") == ["true"]:
    if not st.session_state.recording:
        st.session_state.recording = True
        # 启动录音线程
        threading.Thread(target=start_recording).start()
        st.experimental_set_query_params()  # 清空参数
        st.rerun()

if st.experimental_get_query_params().get("stop_recording") == ["true"]:
    if st.session_state.recording:
        st.session_state.recording = False
        time.sleep(0.5)  # 等待录音线程完成文件保存
        transcribe_audio()
        st.experimental_set_query_params()  # 清空参数
        st.rerun()

# UI展示
st.title("按住空格键录音")
if st.session_state.recording:
    st.warning("正在录音... 松开空格键结束")
else:
    st.info("按住空格键开始说话,松开后自动转文字")

if st.session_state.transcript:
    st.subheader("识别结果:")
    st.write(st.session_state.transcript)

关键说明

  1. 前端按键监听:通过注入JavaScript代码,监听浏览器的keydown和keyup事件,当空格键按下/松开时,通过Streamlit.setComponentValue传递状态给后端。
  2. 线程化录音:录音操作是阻塞的,用threading.Thread启动录音线程,避免卡住Streamlit的UI渲染。
  3. 会话状态管理:用st.session_state跟踪录音状态、临时音频文件路径和识别结果,确保页面刷新时状态不丢失。
  4. 音频转文字:录音停止后,读取临时WAV文件,用speech_recognition调用Google语音识别接口转换文字。

依赖安装

需要安装额外依赖:

pip install streamlit speechrecognition pyaudio

注意:pyaudio在部分系统可能需要手动安装依赖(如Ubuntu需先安装portaudio19-dev)。

内容的提问来源于stack exchange,提问作者GivenX

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 11:38:17