Python SpeechRecognition库无法获取音频,程序卡在Listening状态如何解决?
问题描述
我尝试基于SpeechRecognition 3.8.1库实现语音采集功能已4天,先后查阅了GitHub对应仓库issue、GeeksforGeeks语音助手教程、Stack Overflow多个同类问题帖等大量资料,所有方案均无效。
最初执行sudo apt-get install python-pyaudio python3-pyaudio安装PyAudio失败,经过多次尝试后通过pipwin install pyaudio完成安装。但运行下述代码时,程序始终卡在输出Listening...的环节,无后续响应:
import os import pyttsx3, datetime, pyaudio import speech_recognition as sr # Initial Setup for pyttsx3 - speaking abilities engine = pyttsx3.init("sapi5") voices = engine.getProperty("voices") engine.setProperty("voice", voices[1].id) # 0-male voice , 1-female voice sr.Microphone.list_microphone_names() # Initial Setup for speech_recognition - listening abilities # r.energy_threshold = 10 # print(pyaudio.get_device_count() - 1) def speak(speakable): """speak() takes a string and reads it loud""" engine.say(str(speakable)) engine.runAndWait() def takeCommand(): pyaudio.PyAudio() r = sr.Recognizer() """It takes microphone input from the user and returns string output""" with sr.Microphone() as source: r.adjust_for_ambient_noise(source, duration=0.9) print("Listening...") r.pause_threshold = 45 audio = "" try: audio = r.listen(source) print("Recognizing...") except Exception as e: print("Listen err: ", e) try: print("Recognizing...") query = r.recognize_google(audio) print(f"User said: {query}\n") # User query will be printed. except sr.UnknownValueError as e: print("Say that again please...") return "None" # None string will be returned except Exception as err: print("Check your internet...") return "None" return query def wishMe(): hour = int(datetime.datetime.now().hour) if hour >= 0 and hour < 12: speak("Good Morning!") elif hour >= 12 and hour < 18: speak("Good Afternoon!") else: speak("Good Evening!") speak( "Hello Sir, I am Friday, your Artificial intelligence assistant. Please tell me how may I help you" ) if __name__ == "__main__": os.system("CLS") while True: command = takeCommand().lower() print(f"Command: {command}") if "wish" in command: wishMe()
环境信息
- 运行环境为Windows 10 Home
- 代码编辑器为VS Code
- 本项目未使用virtual env虚拟环境
- 已验证Chrome语音搜索功能可正常使用,麦克风权限无异常
- 执行
python3 -m speech_recognition命令的结果见下图:
解决方案
按顺序修改以下代码问题即可解决卡顿问题:
- 调整pause_threshold参数
你设置的r.pause_threshold = 45代表语音结束后需要等待45秒才会进入识别环节,这是程序看起来卡住的核心原因,将该值调整为1即可,代表语音停顿1秒就判定为输入结束。 - 删除多余的pyaudio初始化代码
每次调用takeCommand()时都执行pyaudio.PyAudio()会重复创建实例,占用麦克风资源,直接删掉这行代码即可,speech_recognition库内部会自动管理pyaudio实例。 - 指定麦克风设备索引
系统存在多个音频输入设备时,默认选中的设备可能不是你正在使用的麦克风。运行print(sr.Microphone.list_microphone_names())输出所有设备列表,找到你在用的麦克风对应的索引值,修改sr.Microphone()为sr.Microphone(device_index=你查到的索引值)即可。 - 修复audio初始值问题
你将audio初始值设为空字符串,如果listen()环节报错,后续识别步骤传入空字符串会触发未知错误,将初始值改为None,识别前增加判空逻辑即可。
修改后的核心takeCommand函数参考:
def takeCommand(): r = sr.Recognizer() # 替换device_index为你查到的麦克风索引 with sr.Microphone(device_index=0) as source: r.adjust_for_ambient_noise(source, duration=0.9) print("Listening...") r.pause_threshold = 1 audio = None try: # 增加超时限制避免程序无限等待 audio = r.listen(source, timeout=10, phrase_time_limit=15) print("Recognizing...") except Exception as e: print("Listen err: ", e) return "None" try: # 如需识别中文可以加language='zh-CN'参数,不需要可删除 query = r.recognize_google(audio) print(f"User said: {query}\n") except sr.UnknownValueError as e: print("Say that again please...") return "None" except Exception as err: print("Check your internet...") return "None" return query
内容的提问来源于stack exchange,提问作者Curious Learner
相关产品推荐
相关产品推荐

