You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:基于Python speech_recognition实现语音助手唤醒词系统

无轮询唤醒词系统实现方案

针对轮询式方案存在识别间隙、漏触发唤醒词的问题,推荐采用实时音频流+本地唤醒词检测模型的方案,核心逻辑是用轻量的唤醒词检测模型持续监听音频,只有检测到唤醒词时才触发完整指令识别,完全避免轮询带来的间隙问题。

这里选用Picovoice的Porcupine实现唤醒词检测,它是本地运行的轻量模型,资源占用低,支持自定义唤醒词,响应速度快,适合语音助手场景。

实现步骤

1. 安装依赖

pip install pvporcupine speechrecognition pyaudio

2. 获取Porcupine Access Key

去Picovoice官方平台免费获取Access Key(用于初始化Porcupine模型)。

3. 整合唤醒词检测与指令识别代码

import pvporcupine
import pyaudio
import speech_recognition as sr

# 配置参数
ACCESS_KEY = "你的Porcupine Access Key"
WAKE_WORD = "jarvis"  # 可选内置唤醒词:alexa, amazon, blueberry, computer, grasshopper, hey google, hey siri, ok google, picovoice, porcupine, terminator

# 初始化Porcupine唤醒词检测器
porcupine = pvporcupine.create(access_key=ACCESS_KEY, keywords=[WAKE_WORD])

# 初始化音频输入
pa = pyaudio.PyAudio()
audio_stream = pa.open(
    rate=porcupine.sample_rate,
    channels=1,
    format=pyaudio.paInt16,
    input=True,
    frames_per_buffer=porcupine.frame_length
)

def takeCommand():
    r = sr.Recognizer()
    with sr.Microphone() as source:
        r.adjust_for_ambient_noise(source)
        print("\nListening for command...")
        r.pause_threshold = 1
        audio = r.listen(source)

    try:
        print("Recognizing command...")
        query = r.recognize_google(audio, language="en-in")
        print(f"Recognized Command: {query}")

    except Exception as e:
        print(e)
        print("I didn't recognize what you said please repeat")
        return "None"

    return query

print(f"Waiting for wake word: 'Hey {WAKE_WORD.title()}'...")

try:
    while True:
        # 读取音频帧进行唤醒词检测
        pcm = audio_stream.read(porcupine.frame_length)
        pcm = pvporcupine.convert_pcm_to_int16(pcm)
        keyword_index = porcupine.process(pcm)

        if keyword_index >= 0:
            print("\nWake word detected!")
            # 触发指令识别
            command = takeCommand()
            # 这里可以添加指令处理逻辑
            if command != "None":
                print(f"Processing command: {command}")
            # 处理完后回到唤醒词监听状态
            print(f"\nWaiting for wake word: 'Hey {WAKE_WORD.title()}'...")
finally:
    # 释放资源
    audio_stream.close()
    pa.terminate()
    porcupine.delete()

方案说明

  • 无间隙监听:Porcupine采用实时音频流处理,持续监听麦克风输入,不会出现轮询方案中识别与监听交替的间隙,确保唤醒词不会漏触发
  • 低资源消耗:Porcupine模型体积小、运行效率高,适合在本地设备持续运行
  • 灵活扩展:支持自定义唤醒词,你可以在Picovoice平台训练自己的专属唤醒词
  • 无缝整合:检测到唤醒词后直接调用你已有的takeCommand函数,无需大幅修改原有代码

内容的提问来源于stack exchange,提问作者Aditya Chandra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 03:50:20