You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中transcription函数执行顺序异常:API先于提示音执行求助

问题:语音识别函数执行顺序异常,提示音晚于API调用

我的transcription函数本应先播放提示音,告知用户可以对着麦克风说话,但实际却先执行了Google语音识别API,正常的线性执行顺序被打破。相关代码如下:

r = sr.Recognizer()
### instantiate pycharm
p = pyaudio.PyAudio()


acknowledge_wav = ["asyouwish.wav", "openingdesired.wav", "rightaway.wav"]

keywords = {"firefox": "C:/Program Files (x86)/Mozilla Firefox/firefox.exe",
            "discord": "C:/Users/LucasGames/AppData/Local/Discord/Update.exe --processStart Discord.exe",
            "pycharm": "C:/Program Files/JetBrains/PyCharm Community Edition 2022.1/bin/pycharm64.exe",
            "chrome": "C:/Program Files/Google/Chrome/Application/chrome.exe",
            "wars": "C:/Spiele/steamapps/common/Jedi Outcast/GameData/JK2MV/jk2mvmp.exe"
            }

with sr.Microphone(sample_rate=20000) as mic:
    r.adjust_for_ambient_noise(mic, duration=0.5)
    audio = r.listen(mic)

### play some sound
def play_sound(file):
    wf = wave.open(file)
    stream = p.open(format=p.get_format_from_width(wf.getsampwidth()),
                    channels=wf.getnchannels(),
                    rate=wf.getframerate(),
                    output=True)
    wav_data = wf.readframes(1024)

    while len(wav_data) > 0:
        stream.write(wav_data)
        wav_data = wf.readframes(1024)
    stream.stop_stream()
    stream.close()
    p.terminate()

### activation and transcription of recognizer
def transcription(audio_data):
    ## speak now indicator
    play_sound("notificationtest.wav")
 
    ## call google API
    try:
        command = r.recognize_google(audio_data, language="en-USA").lower()
    except sr.RequestError:
        print("I am experiencing brain fog - please try again.")
    except sr.UnknownValueError:
        print("I could not understand you. Please repeat your command.")

    # splits the command sentence in single words to iterate over
    splt_command = command.split()

    ## check if keyword is in the spoken command
    for word in splt_command:
        if word in keywords:
            print(f"Command: {command}")
            print(f"Keyword: {word}")
            assistant_response = random.choice(acknowledge_wav)
            play_sound(assistant_response)
            subprocess.call(keywords[word])

    # no keyword in command
    print("No keyword detected.")
    print(f"Said: {command}")

transcription(audio)

原因分析

问题出在代码执行顺序上:

  • 在调用transcription(audio)之前,代码已经执行了with sr.Microphone()块中的r.listen(mic),也就是麦克风录音操作在transcription函数执行前就已经完成。
  • 当进入transcription函数时,先播放提示音,但此时录音早就结束了,之后调用recognize_google只是识别已经录好的音频数据,所以看起来像是API调用先于提示音播放。

修复方案

  1. 调整录音逻辑位置:将麦克风录音的代码移到transcription函数内部,放在提示音播放之后,确保流程是:播放提示音 → 开始录音 → 调用识别API。
  2. 修复PyAudio实例重复终止问题:play_sound函数中的p.terminate()会直接终止PyAudio实例,导致后续再次调用play_sound时出错,需要将该操作移到程序结束时执行。
  3. 处理变量未定义情况:当语音识别失败时,command变量可能未定义,需要添加默认值或在异常处理中提前返回,避免后续command.split()报错。

修改后的完整代码

import speech_recognition as sr
import pyaudio
import wave
import subprocess
import random

r = sr.Recognizer()
p = pyaudio.PyAudio()

acknowledge_wav = ["asyouwish.wav", "openingdesired.wav", "rightaway.wav"]

keywords = {"firefox": "C:/Program Files (x86)/Mozilla Firefox/firefox.exe",
            "discord": "C:/Users/LucasGames/AppData/Local/Discord/Update.exe --processStart Discord.exe",
            "pycharm": "C:/Program Files/JetBrains/PyCharm Community Edition 2022.1/bin/pycharm64.exe",
            "chrome": "C:/Program Files/Google/Chrome/Application/chrome.exe",
            "wars": "C:/Spiele/steamapps/common/Jedi Outcast/GameData/JK2MV/jk2mvmp.exe"
            }

def play_sound(file):
    wf = wave.open(file, 'rb')
    stream = p.open(format=p.get_format_from_width(wf.getsampwidth()),
                    channels=wf.getnchannels(),
                    rate=wf.getframerate(),
                    output=True)
    wav_data = wf.readframes(1024)

    while len(wav_data) > 0:
        stream.write(wav_data)
        wav_data = wf.readframes(1024)
    
    stream.stop_stream()
    stream.close()
    wf.close()

def transcription():
    # 先播放提示音
    play_sound("notificationtest.wav")
    
    # 提示音播放后开始录音
    with sr.Microphone(sample_rate=20000) as mic:
        r.adjust_for_ambient_noise(mic, duration=0.5)
        print("Listening...")
        audio = r.listen(mic)
    
    # 调用Google语音识别API
    command = None
    try:
        command = r.recognize_google(audio, language="en-USA").lower()
        print(f"Recognized: {command}")
    except sr.RequestError:
        print("I am experiencing brain fog - please try again.")
        p.terminate()
        return
    except sr.UnknownValueError:
        print("I could not understand you. Please repeat your command.")
        p.terminate()
        return

    # 处理识别结果
    splt_command = command.split()
    keyword_found = False
    
    for word in splt_command:
        if word in keywords:
            keyword_found = True
            print(f"Command: {command}")
            print(f"Keyword: {word}")
            assistant_response = random.choice(acknowledge_wav)
            play_sound(assistant_response)
            subprocess.call(keywords[word])
            break
    
    if not keyword_found:
        print("No keyword detected.")
        print(f"Said: {command}")
    
    # 程序结束时终止PyAudio实例
    p.terminate()

# 调用transcription函数
transcription()

内容的提问来源于stack exchange,提问作者Lucas Szikszai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 04:36:17