使用SpeechRecognition开发类Alexa机器人时麦克风无法拾音
问题说明
开发类Alexa语音机器人过程中,使用SpeechRecognition模块实现语音输入功能,原有代码如下:
from datetime import datetime import pyttsx3 import speech_recognition as sr engine = pyttsx3.init('sapi5') voices = engine.getProperty('voices') # print(voices) engine.setProperty('voice', voices[1].id) def speak(audio): '''It speak out the audio to the user''' engine.say(audio) print(f"Bot: {audio}") engine.runAndWait() def wishMe(): '''It wish the user through datetime module''' hour = int(datetime.now().hour) if hour >= 0 and hour < 12: speak("Good Morning!") elif hour >= 12 and hour < 18: speak("Good Afternoon!") else: speak("Good Evening!") speak("I am AI bot. Please tell me how can I help you.") def takeCommand(): '''It takes command from the users's microphone''' r = sr.Recognizer() with sr.Microphone() as source: print("Listening . . .") audio = r.listen(source) try: print("Recoginizing . . .") query = r.recognize_google(audio, language='en_in') print(f"You: {query}\n") except Exception as e: print(e) speak("Say that again please") return "None" return query if __name__ == '__main__': wishMe() takeCommand()
运行代码时程序无法拾取麦克风输入的语音,需在尽量不大幅修改原有代码的前提下解决问题。
排查步骤与修复方案
所有修改均不改动原有业务逻辑,仅补充必要配置:
- 补充环境噪声校准:原有代码直接启动监听,没有根据当前环境底噪调整识别阈值,是最常见的拾音失败原因。在监听前加1秒环境噪声采样即可,不需要改其他逻辑。
- 确认麦克风设备正确:多数电脑存在多个音频输入设备(摄像头麦、耳机麦、虚拟音频设备等),
SpeechRecognition默认选中的设备不一定是你正在用的麦。先运行以下代码打印所有可用麦克风列表,找到你在用的设备对应的序号:
import speech_recognition as sr print(sr.Microphone.list_microphone_names())
- 补充监听超时配置:默认监听参数没有等待时长限制,容易出现程序假死、看起来没在拾音的情况,给
listen方法加超时和单句时长限制即可。 - 检查系统权限:Windows/macOS需要给运行代码的IDE/终端开放麦克风权限,否则程序无法访问音频输入设备。
- 修复依赖问题:Windows环境下直接通过
pip install pyaudio经常会出现安装不全、运行异常的问题,用pipwin安装预编译版本即可:
pip install pipwin pipwin install pyaudio
修改后可直接运行的核心代码
仅对takeCommand函数做了最小改动,其余代码完全保留:
def takeCommand(): '''It takes command from the users's microphone''' r = sr.Recognizer() # 把device_index的值替换成你自己查到的麦克风序号即可,单设备可删除该参数 with sr.Microphone(device_index=1) as source: # 新增:环境噪声校准,1秒采样自动调整识别阈值 r.adjust_for_ambient_noise(source, duration=1) print("Listening . . .") # 新增:超时配置,5秒内没检测到语音就抛出异常,单句最长采集10秒 audio = r.listen(source, timeout=5, phrase_time_limit=10) try: print("Recognizing . . .") # 修正原代码的拼写错误 query = r.recognize_google(audio, language='en_in') print(f"You: {query}\n") except Exception as e: print(e) speak("Say that again please") return "None" return query
内容的提问来源于stack exchange,提问作者myNameisShourya
相关产品推荐
相关产品推荐

