Python中transcription函数执行顺序异常:API先于提示音执行求助
问题:语音识别函数执行顺序异常,提示音晚于API调用
我的transcription函数本应先播放提示音,告知用户可以对着麦克风说话,但实际却先执行了Google语音识别API,正常的线性执行顺序被打破。相关代码如下:
r = sr.Recognizer() ### instantiate pycharm p = pyaudio.PyAudio() acknowledge_wav = ["asyouwish.wav", "openingdesired.wav", "rightaway.wav"] keywords = {"firefox": "C:/Program Files (x86)/Mozilla Firefox/firefox.exe", "discord": "C:/Users/LucasGames/AppData/Local/Discord/Update.exe --processStart Discord.exe", "pycharm": "C:/Program Files/JetBrains/PyCharm Community Edition 2022.1/bin/pycharm64.exe", "chrome": "C:/Program Files/Google/Chrome/Application/chrome.exe", "wars": "C:/Spiele/steamapps/common/Jedi Outcast/GameData/JK2MV/jk2mvmp.exe" } with sr.Microphone(sample_rate=20000) as mic: r.adjust_for_ambient_noise(mic, duration=0.5) audio = r.listen(mic) ### play some sound def play_sound(file): wf = wave.open(file) stream = p.open(format=p.get_format_from_width(wf.getsampwidth()), channels=wf.getnchannels(), rate=wf.getframerate(), output=True) wav_data = wf.readframes(1024) while len(wav_data) > 0: stream.write(wav_data) wav_data = wf.readframes(1024) stream.stop_stream() stream.close() p.terminate() ### activation and transcription of recognizer def transcription(audio_data): ## speak now indicator play_sound("notificationtest.wav") ## call google API try: command = r.recognize_google(audio_data, language="en-USA").lower() except sr.RequestError: print("I am experiencing brain fog - please try again.") except sr.UnknownValueError: print("I could not understand you. Please repeat your command.") # splits the command sentence in single words to iterate over splt_command = command.split() ## check if keyword is in the spoken command for word in splt_command: if word in keywords: print(f"Command: {command}") print(f"Keyword: {word}") assistant_response = random.choice(acknowledge_wav) play_sound(assistant_response) subprocess.call(keywords[word]) # no keyword in command print("No keyword detected.") print(f"Said: {command}") transcription(audio)
原因分析
问题出在代码执行顺序上:
- 在调用
transcription(audio)之前,代码已经执行了with sr.Microphone()块中的r.listen(mic),也就是麦克风录音操作在transcription函数执行前就已经完成。 - 当进入transcription函数时,先播放提示音,但此时录音早就结束了,之后调用
recognize_google只是识别已经录好的音频数据,所以看起来像是API调用先于提示音播放。
修复方案
- 调整录音逻辑位置:将麦克风录音的代码移到transcription函数内部,放在提示音播放之后,确保流程是:播放提示音 → 开始录音 → 调用识别API。
- 修复PyAudio实例重复终止问题:
play_sound函数中的p.terminate()会直接终止PyAudio实例,导致后续再次调用play_sound时出错,需要将该操作移到程序结束时执行。 - 处理变量未定义情况:当语音识别失败时,
command变量可能未定义,需要添加默认值或在异常处理中提前返回,避免后续command.split()报错。
修改后的完整代码
import speech_recognition as sr import pyaudio import wave import subprocess import random r = sr.Recognizer() p = pyaudio.PyAudio() acknowledge_wav = ["asyouwish.wav", "openingdesired.wav", "rightaway.wav"] keywords = {"firefox": "C:/Program Files (x86)/Mozilla Firefox/firefox.exe", "discord": "C:/Users/LucasGames/AppData/Local/Discord/Update.exe --processStart Discord.exe", "pycharm": "C:/Program Files/JetBrains/PyCharm Community Edition 2022.1/bin/pycharm64.exe", "chrome": "C:/Program Files/Google/Chrome/Application/chrome.exe", "wars": "C:/Spiele/steamapps/common/Jedi Outcast/GameData/JK2MV/jk2mvmp.exe" } def play_sound(file): wf = wave.open(file, 'rb') stream = p.open(format=p.get_format_from_width(wf.getsampwidth()), channels=wf.getnchannels(), rate=wf.getframerate(), output=True) wav_data = wf.readframes(1024) while len(wav_data) > 0: stream.write(wav_data) wav_data = wf.readframes(1024) stream.stop_stream() stream.close() wf.close() def transcription(): # 先播放提示音 play_sound("notificationtest.wav") # 提示音播放后开始录音 with sr.Microphone(sample_rate=20000) as mic: r.adjust_for_ambient_noise(mic, duration=0.5) print("Listening...") audio = r.listen(mic) # 调用Google语音识别API command = None try: command = r.recognize_google(audio, language="en-USA").lower() print(f"Recognized: {command}") except sr.RequestError: print("I am experiencing brain fog - please try again.") p.terminate() return except sr.UnknownValueError: print("I could not understand you. Please repeat your command.") p.terminate() return # 处理识别结果 splt_command = command.split() keyword_found = False for word in splt_command: if word in keywords: keyword_found = True print(f"Command: {command}") print(f"Keyword: {word}") assistant_response = random.choice(acknowledge_wav) play_sound(assistant_response) subprocess.call(keywords[word]) break if not keyword_found: print("No keyword detected.") print(f"Said: {command}") # 程序结束时终止PyAudio实例 p.terminate() # 调用transcription函数 transcription()
内容的提问来源于stack exchange,提问作者Lucas Szikszai
相关产品推荐
相关产品推荐

