Python语音识别助手播放MP3时如何屏蔽自识别并实现唤醒触发
问题
我开发了一款Python语音识别助手,它会播放下载的MP3音频,且已将MP3播放逻辑放在后台独立线程中。目前存在的问题是,语音识别会检测MP3音频的内容并作出响应,我希望让语音识别保持静默,直到我说出特定唤醒语音才触发。
现有代码
播放与获取MP3的函数
def play_quran(): speak("Ready to play Quran. Tell me which Surah number you want to hear.") #qari_num = input("Enter Surah Number: ") qari_num = recordAudio() url = ("https://api.quran.com/api/v4/chapter_recitations/9/" + str(qari_num)) print(url) response = requests.get(url) my_dictionary = requests.get(url).json() rdata = response.json() print(json.dumps(my_dictionary, indent=4)) surah_to_play = (my_dictionary['audio_file']['audio_url']) print(surah_to_play) response = request.urlretrieve(surah_to_play, qari_num + ".mp3") os.system("mpg123 -q " + qari_num + ".mp3") stop_listening = sr.Recognizer().listen_in_background(sr.Microphone(), recordAudio) # time.sleep(2) # exit()
函数调用代码
if "play Quran" in data: speak("opening Quran. One moment please") t = threading.Thread( target=play_quran) # < Note that I did not actually call the function, but instead sent it as a parameter t.daemon = True t.start() # < This actually starts the thread execution in the background
解决方法
核心是在MP3播放期间暂停语音识别响应,仅在播放结束后恢复,且恢复后只响应特定唤醒词,具体实现如下:
1. 全局控制监听状态
添加线程安全的全局变量,控制语音识别是否处理输入:
import threading # 控制是否激活语音识别响应 is_listening_active = True # 线程锁,避免多线程操作全局变量冲突 lock = threading.Lock()
2. 修改语音识别函数(recordAudio)
让函数只在激活状态下工作,且仅响应特定唤醒词:
def recordAudio(): global is_listening_active # 先检查是否允许响应 with lock: if not is_listening_active: return "" # 原录音识别逻辑 r = sr.Recognizer() with sr.Microphone() as source: audio = r.listen(source) try: data = r.recognize_google(audio) # 替换成你的唤醒词,比如"唤醒助手" if "唤醒助手" not in data: return "" # 非唤醒词,直接忽略 return data except sr.UnknownValueError: return "" except sr.RequestError as e: return ""
3. 调整播放函数(play_quran)
播放MP3前后切换监听状态,同时优化重复请求的问题:
def play_quran(): global is_listening_active speak("Ready to play Quran. Tell me which Surah number you want to hear.") qari_num = recordAudio() # 如果没检测到唤醒词,直接退出 if not qari_num: return # 优化:只请求一次API url = f"https://api.quran.com/api/v4/chapter_recitations/9/{qari_num}" print(url) response = requests.get(url) my_dictionary = response.json() print(json.dumps(my_dictionary, indent=4)) surah_to_play = my_dictionary['audio_file']['audio_url'] print(surah_to_play) # 修正笔误:request → requests requests.urlretrieve(surah_to_play, f"{qari_num}.mp3") # 播放前暂停语音识别响应 with lock: is_listening_active = False # 播放音频 os.system(f"mpg123 -q {qari_num}.mp3") # 播放结束恢复响应 with lock: is_listening_active = True # 重新启动后台监听 stop_listening = sr.Recognizer().listen_in_background(sr.Microphone(), recordAudio)
关键说明
- 线程锁
lock确保全局变量is_listening_active在多线程环境下不会出现冲突; - 播放MP3时关闭识别响应,避免MP3内容被误识别;
- 恢复监听后,只有说出设定的唤醒词才会触发后续操作,避免无效响应;
- 优化了原代码中重复调用
requests.get(url)的问题,减少API请求次数。
内容的提问来源于stack exchange,提问作者ironmantis7x
相关产品推荐
相关产品推荐

