You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python语音识别助手播放MP3时如何屏蔽自识别并实现唤醒触发

问题

我开发了一款Python语音识别助手,它会播放下载的MP3音频,且已将MP3播放逻辑放在后台独立线程中。目前存在的问题是,语音识别会检测MP3音频的内容并作出响应,我希望让语音识别保持静默,直到我说出特定唤醒语音才触发。

现有代码

播放与获取MP3的函数

def play_quran():
    speak("Ready to play Quran. Tell me which Surah number you want to hear.")
    #qari_num = input("Enter Surah Number: ")
    qari_num = recordAudio()
    url = ("https://api.quran.com/api/v4/chapter_recitations/9/" + str(qari_num))
    print(url)
    response = requests.get(url)
    my_dictionary = requests.get(url).json()
    rdata = response.json()
    print(json.dumps(my_dictionary, indent=4))
    surah_to_play = (my_dictionary['audio_file']['audio_url'])
    print(surah_to_play)
    response = request.urlretrieve(surah_to_play, qari_num + ".mp3")
    os.system("mpg123 -q " + qari_num + ".mp3")
    stop_listening = sr.Recognizer().listen_in_background(sr.Microphone(), recordAudio)
#    time.sleep(2)
#    exit()

函数调用代码

if "play Quran" in data:
    speak("opening Quran. One moment please")
    t = threading.Thread(
        target=play_quran)  # < Note that I did not actually call the function, but instead sent it as a parameter
    t.daemon = True
    t.start()  # < This actually starts the thread execution in the background
解决方法

核心是在MP3播放期间暂停语音识别响应,仅在播放结束后恢复,且恢复后只响应特定唤醒词,具体实现如下:

1. 全局控制监听状态

添加线程安全的全局变量,控制语音识别是否处理输入:

import threading
# 控制是否激活语音识别响应
is_listening_active = True
# 线程锁,避免多线程操作全局变量冲突
lock = threading.Lock()

2. 修改语音识别函数(recordAudio)

让函数只在激活状态下工作,且仅响应特定唤醒词:

def recordAudio():
    global is_listening_active
    # 先检查是否允许响应
    with lock:
        if not is_listening_active:
            return ""
    
    # 原录音识别逻辑
    r = sr.Recognizer()
    with sr.Microphone() as source:
        audio = r.listen(source)
    
    try:
        data = r.recognize_google(audio)
        # 替换成你的唤醒词,比如"唤醒助手"
        if "唤醒助手" not in data:
            return ""  # 非唤醒词,直接忽略
        return data
    except sr.UnknownValueError:
        return ""
    except sr.RequestError as e:
        return ""

3. 调整播放函数(play_quran)

播放MP3前后切换监听状态,同时优化重复请求的问题:

def play_quran():
    global is_listening_active
    speak("Ready to play Quran. Tell me which Surah number you want to hear.")
    qari_num = recordAudio()
    # 如果没检测到唤醒词,直接退出
    if not qari_num:
        return
    
    # 优化:只请求一次API
    url = f"https://api.quran.com/api/v4/chapter_recitations/9/{qari_num}"
    print(url)
    response = requests.get(url)
    my_dictionary = response.json()
    print(json.dumps(my_dictionary, indent=4))
    
    surah_to_play = my_dictionary['audio_file']['audio_url']
    print(surah_to_play)
    # 修正笔误:request → requests
    requests.urlretrieve(surah_to_play, f"{qari_num}.mp3")
    
    # 播放前暂停语音识别响应
    with lock:
        is_listening_active = False
    # 播放音频
    os.system(f"mpg123 -q {qari_num}.mp3")
    # 播放结束恢复响应
    with lock:
        is_listening_active = True
    
    # 重新启动后台监听
    stop_listening = sr.Recognizer().listen_in_background(sr.Microphone(), recordAudio)

关键说明

  • 线程锁lock确保全局变量is_listening_active在多线程环境下不会出现冲突;
  • 播放MP3时关闭识别响应,避免MP3内容被误识别;
  • 恢复监听后,只有说出设定的唤醒词才会触发后续操作,避免无效响应;
  • 优化了原代码中重复调用requests.get(url)的问题,减少API请求次数。

内容的提问来源于stack exchange,提问作者ironmantis7x

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 15:33:26