You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用SpeechRecognition开发类Alexa机器人时麦克风无法拾音

问题说明

开发类Alexa语音机器人过程中,使用SpeechRecognition模块实现语音输入功能,原有代码如下:

from datetime import datetime
import pyttsx3
import speech_recognition as sr

engine = pyttsx3.init('sapi5')
voices = engine.getProperty('voices')
# print(voices)
engine.setProperty('voice', voices[1].id)

def speak(audio):
    '''It speak out the audio to the user'''
    engine.say(audio)
    print(f"Bot: {audio}")
    engine.runAndWait()

def wishMe():
    '''It wish the user through datetime module'''
    hour = int(datetime.now().hour)
    if hour >= 0 and hour < 12:
        speak("Good Morning!")
    elif hour >= 12 and hour < 18:
        speak("Good Afternoon!")
    else:
        speak("Good Evening!")

    speak("I am AI bot. Please tell me how can I help you.")

def takeCommand():
    '''It takes command from the users's microphone'''
    r = sr.Recognizer()
    with sr.Microphone() as source:
        print("Listening . . .")
        audio = r.listen(source)
 
    try:
        print("Recoginizing . . .")
        query = r.recognize_google(audio, language='en_in')
        print(f"You: {query}\n")
    except Exception as e:
        print(e)
        speak("Say that again please")
        return "None"
    return query



if __name__ == '__main__':
    wishMe()
    takeCommand()

运行代码时程序无法拾取麦克风输入的语音,需在尽量不大幅修改原有代码的前提下解决问题。

排查步骤与修复方案

所有修改均不改动原有业务逻辑,仅补充必要配置:

  • 补充环境噪声校准:原有代码直接启动监听,没有根据当前环境底噪调整识别阈值,是最常见的拾音失败原因。在监听前加1秒环境噪声采样即可,不需要改其他逻辑。
  • 确认麦克风设备正确:多数电脑存在多个音频输入设备(摄像头麦、耳机麦、虚拟音频设备等),SpeechRecognition默认选中的设备不一定是你正在用的麦。先运行以下代码打印所有可用麦克风列表,找到你在用的设备对应的序号:
import speech_recognition as sr
print(sr.Microphone.list_microphone_names())
  • 补充监听超时配置:默认监听参数没有等待时长限制,容易出现程序假死、看起来没在拾音的情况,给listen方法加超时和单句时长限制即可。
  • 检查系统权限:Windows/macOS需要给运行代码的IDE/终端开放麦克风权限,否则程序无法访问音频输入设备。
  • 修复依赖问题:Windows环境下直接通过pip install pyaudio经常会出现安装不全、运行异常的问题,用pipwin安装预编译版本即可:
pip install pipwin
pipwin install pyaudio
修改后可直接运行的核心代码

仅对takeCommand函数做了最小改动,其余代码完全保留:

def takeCommand():
    '''It takes command from the users's microphone'''
    r = sr.Recognizer()
    # 把device_index的值替换成你自己查到的麦克风序号即可,单设备可删除该参数
    with sr.Microphone(device_index=1) as source:
        # 新增:环境噪声校准,1秒采样自动调整识别阈值
        r.adjust_for_ambient_noise(source, duration=1)
        print("Listening . . .")
        # 新增:超时配置,5秒内没检测到语音就抛出异常,单句最长采集10秒
        audio = r.listen(source, timeout=5, phrase_time_limit=10)
 
    try:
        print("Recognizing . . .") # 修正原代码的拼写错误
        query = r.recognize_google(audio, language='en_in')
        print(f"You: {query}\n")
    except Exception as e:
        print(e)
        speak("Say that again please")
        return "None"
    return query

内容的提问来源于stack exchange,提问作者myNameisShourya

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 03:33:32