You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python虚拟助手优化:实现等待用户发声后再启动语音监听

解决Python虚拟助手持续监听问题:仅在用户发声时启动监听

我完全懂你的困扰——当前助手一直卡着监听状态,没法做到只在用户开口说话时才启动识别。原代码里的listen(mic)是阻塞式的,会一直死等音频输入,所以程序始终停在监听环节。下面是调整后的完整实现,通过音频能量检测来实现“等用户发声后再启动监听”的逻辑:

修改后的可运行代码

import speech_recognition
import pyttsx3 as tts
import subprocess
import datetime
import webbrowser
import time

recognizer = speech_recognition.Recognizer()
speaker = tts.init()
# 配置语音和语速
voices = speaker.getProperty('voices')
speaker.setProperty('voice', voices[1].id)
speaker.setProperty('rate', 175)

def detect_voice_activity(mic):
    """检测用户是否发声:音频能量超过环境噪音阈值时触发"""
    # 先校准环境噪音,设置基准阈值
    recognizer.adjust_for_ambient_noise(mic, duration=0.5)
    print("等待用户发声...")
    while True:
        # 短时间监听音频,判断能量值
        audio = recognizer.listen(mic, timeout=None, phrase_time_limit=1)
        # 当音频能量超过阈值,判定用户开始说话
        if recognizer.energy_threshold < audio.energy:
            return True

def record_audio(ask=False):
    with speech_recognition.Microphone() as mic:
        if ask:
            speaker.speak(ask)
        
        # 先等待用户发声,再启动正式识别
        detect_voice_activity(mic)
        
        print("正在监听你的指令...")
        voice_data = ''
        recognizer.adjust_for_ambient_noise(mic, 0.05)
        # 设置短语时间限制,避免无意义的长时间监听
        audio = recognizer.listen(mic, phrase_time_limit=5)
        try:
            voice_data = recognizer.recognize_google(audio, language="en-IN")
            print(f"识别到指令:{voice_data}")
        except speech_recognition.UnknownValueError:
            speaker.speak('Sorry, I did not understand what you just said. Please try again.')
        except speech_recognition.RequestError:
            speaker.speak("Sorry, my speech service is down for the time being. Please try again later.")
        return voice_data

def responses(command):
    if not command:
        return
    if 'hello' in command:
        speaker.speak("hello sir, how can I help you.")
    elif 'what is your name' in command:
        speaker.speak("My name is Otto Octavius")
    elif 'time' in command:
        current_time = datetime.datetime.now().strftime("%I:%M:%S")
        speaker.speak(current_time)
    elif 'date' in command:
        current_date = datetime.datetime.now().strftime("%Y-%m-%d")
        speaker.speak(current_date)
    elif 'open' and 'telegram' in command:
        speaker.speak("opening telegram")
        subprocess.Popen("D:\\My Folder\\My Softwares\\Telegram Desktop\\Telegram.exe")
    elif 'close' and 'telegram' and 'window' in command:
        speaker.speak("closing telegram")
        subprocess.call(["taskkill","/F","/IM","Telegram.exe"])
    elif 'open' and 'binance' in command:
        speaker.speak("opening binance")
        subprocess.Popen("D:\\My Folder\\My Softwares\\Binance\\Binance.exe")
    elif 'close' and 'binance' in command:
        speaker.speak('closing binance')
        subprocess.call(["taskkill" , "/F" , "/IM" , 'Binance.exe'])
    elif 'my folder' in command:
        speaker.speak('opening my folder')
        webbrowser.open("D:\\My Folder")
    elif 'search' in command:
        search_object = record_audio("What do you want me to search for?")
        if search_object:
            url = f"https://www.google.com/search?q={search_object}"
            speaker.speak(f'Searching for {search_object}')
            webbrowser.get('C:/Program Files (x86)/Google/Chrome/Application/chrome.exe %s').open_new_tab(url)
    elif any(phrase in command for phrase in ['bye', 'good bye', 'good night']):
        speaker.speak('Good bye sir, have a nice day ahead' if 'night' not in command else "Good Night")
        exit()

time.sleep(1)
speaker.speak("Welcome, how can I help you")
while True:
    command = record_audio()
    responses(command)

核心改进说明

  1. 新增语音活动检测函数:detect_voice_activity会持续监听麦克风输入,只有当音频能量超过环境噪音阈值时,才会触发后续的正式识别,解决了持续监听的问题。
  2. 修复原代码隐性bug:比如原代码里tts.speak是错误调用(应该用初始化后的speaker对象)、close binance命令漏了/IM参数、时间日期始终显示初始化时的固定值等问题。
  3. 优化监听效率:添加phrase_time_limit=5,避免程序长时间无意义监听,提升整体响应速度。

这样调整后,你的虚拟助手会先处于“待机等待”状态,只有当用户开口说话时才会启动监听和识别逻辑。

内容的提问来源于stack exchange,提问作者pip install logic

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:49:57