You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中实现实时自动语音识别(监听用户说话)?

高效实现麦克风的按键说话与自动语音识别模式

现有代码的核心问题

  1. 每次循环重复初始化Recognizer和Microphone对象,造成资源浪费
  2. listen()默认的静音超时机制(约8秒),导致中途说话无法被及时捕捉,只能等超时触发异常后才进入下一轮监听
  3. 异常处理过于笼统,无法区分「未识别到语音」和「识别服务故障」的不同情况

一、自动语音识别模式优化实现

通过校准环境噪音、动态调整阈值、复用资源的方式,实现「检测到说话才触发识别」的自动模式:

import speech_recognition as sr

# 全局初始化识别器和麦克风,避免重复创建资源
recognizer = sr.Recognizer()
# 开启动态能量阈值,自动适应环境噪音变化
recognizer.dynamic_energy_threshold = True
microphone = sr.Microphone()

def auto_recognition():
    with microphone as source:
        # 启动时校准环境噪音,仅需执行一次(环境变化大时可定期重新校准)
        print("正在校准环境噪音,请保持安静...")
        recognizer.adjust_for_ambient_noise(source, duration=1)
        print("校准完成,等待您说话...")
        
        while True:
            try:
                # 持续监听,直到检测到说话并静音结束,无超时限制
                audio = recognizer.listen(source, timeout=None)
                # 调用谷歌语音识别(指定中文语言)
                text = recognizer.recognize_google(audio, language="zh-CN")
                print(f"您说的是:{text}")
            except sr.UnknownValueError:
                # 未识别到有效语音,继续监听
                continue
            except sr.RequestError as e:
                # 识别服务故障,终止程序
                print(f"语音识别服务出错:{e}")
                break

if __name__ == "__main__":
    auto_recognition()

优化点说明

  • 复用Recognizer和Microphone对象,减少资源开销
  • dynamic_energy_threshold = True让识别器自动调整噪音阈值,适配不同环境
  • adjust_for_ambient_noise校准背景音,确保准确区分噪音和说话声
  • listen(timeout=None)取消超时限制,仅在检测到完整语音片段(说话后静音)时返回

二、按键说话模式实现

借助pynput库监听按键事件,实现「按住按键录制、松开识别」的模式(需先执行pip install pynput安装依赖):

import speech_recognition as sr
from pynput.keyboard import Listener, Key

recognizer = sr.Recognizer()
microphone = sr.Microphone()
is_recording = False
audio_data = None

def on_press(key):
    global is_recording, audio_data
    # 按住空格键开始录制
    if key == Key.space and not is_recording:
        print("开始录制,请说话...")
        is_recording = True
        with microphone as source:
            # 快速校准环境噪音
            recognizer.adjust_for_ambient_noise(source, duration=0.5)
            # 持续录制直到按键松开
            audio_data = recognizer.listen(source, timeout=None)

def on_release(key):
    global is_recording, audio_data
    # 松开空格键结束录制并识别
    if key == Key.space and is_recording:
        is_recording = False
        print("录制结束,正在识别...")
        try:
            text = recognizer.recognize_google(audio_data, language="zh-CN")
            print(f"您说的是:{text}")
        except sr.UnknownValueError:
            print("抱歉,没识别到您说的内容")
        except sr.RequestError as e:
            print(f"语音识别服务出错:{e}")
    # 按ESC键退出程序
    elif key == Key.esc:
        return False

def push_to_talk():
    print("按键说话模式:按住空格键说话,松开后识别;按ESC键退出")
    with Listener(on_press=on_press, on_release=on_release) as listener:
        listener.join()

if __name__ == "__main__":
    push_to_talk()

功能说明

  • 按住空格键触发录制,松开后立即识别语音
  • 支持ESC键快速退出程序
  • 同样复用核心对象,避免重复初始化

内容的提问来源于stack exchange,提问作者shree varshan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 16:39:21