You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现离线持续运行的语音识别服务

Offline Continuous Speech Recognition with Targeted Command Response

Hey there! Let's work through your problem—building an offline, continuously running speech recognition service that only reacts to specific commands, not every random word you say. I’ve dealt with similar headaches with CMU Sphinx before, so here’s what I recommend:

Fixing the CMU Sphinx "Persistent Hypothesis" Issue

If you want to stick with CMU Sphinx, the key is to constrain what it listens for and reset its state after each command:

  • Use a Custom Grammar (JSGF)
    Ditch the generic language model—instead, define a strict grammar that only includes your target commands. For example, create a commands.jsgf file like this:

    grammar commands;
    public <command> = turn on living room light | turn off kitchen light | set thermostat to 22 degrees;
    

    When you initialize the Sphinx decoder, load this grammar instead of the default LM. This tells the engine to only look for those exact command phrases, not arbitrary speech.

  • Manually Reset the Decoder State
    Sphinx holds onto previous hypotheses by default. After you process a recognized command, call decoder.end_utt() to close the current utterance, then decoder.start_utt() to start fresh. This clears the old state so it doesn’t carry over partial matches to the next listening cycle.

  • Add a Wake Word Trigger (Optional)
    For better control, set up a wake word (like "Hey Assistant") that activates command listening. Use a lightweight wake word model first—only when the wake word is detected do you switch to your command grammar. This prevents accidental triggers from random speech.

Better Alternatives for Offline Continuous Recognition

If CMU Sphinx feels clunky, these tools are more streamlined for offline, targeted recognition:

Vosk (Based on Kaldi)

Vosk is designed for offline real-time ASR and makes it trivial to filter for specific commands. It has pre-trained lightweight models, and you can even pass a list of keywords to prioritize:

import vosk
import pyaudio
import json

# Load a pre-trained small offline model (download locally first)
model = vosk.Model("model-small")
recognizer = vosk.KaldiRecognizer(model, 16000)

# Initialize microphone input
p = pyaudio.PyAudio()
stream = p.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True, frames_per_buffer=8192)
stream.start_stream()

print("Listening for commands...")

while True:
    data = stream.read(4096)
    if recognizer.AcceptWaveform(data):
        result = json.loads(recognizer.Result())
        command = result["text"].lower()
        # Check for your target commands
        if "turn on light" in command:
            print("✅ Executing: Turn on light")
        elif "turn off light" in command:
            print("✅ Executing: Turn off light")
        # Add more command checks here

stream.stop_stream()
stream.close()
p.terminate()

OpenAI Whisper

Whisper’s tiny/base models work offline and are great for continuous recognition. You can run real-time audio capture and filter results to only trigger on your commands:

import whisper
import pyaudio
import numpy as np

# Load lightweight offline model
model = whisper.load_model("tiny")

# Audio setup
p = pyaudio.PyAudio()
stream = p.open(format=pyaudio.paInt16, channels=1, rate=16000, input=True, frames_per_buffer=1024)

print("Listening...")

while True:
    # Capture audio chunk
    audio_data = np.frombuffer(stream.read(4096), dtype=np.int16)
    # Convert to float32 (Whisper's expected format)
    audio_float = audio_data.astype(np.float32) / 32768.0
    # Transcribe offline
    result = model.transcribe(audio_float, fp16=False)
    text = result["text"].lower()
    # Trigger only on target commands
    if "set temperature to 25" in text:
        print("✅ Executing: Set temperature to 25°C")

Key Tips for Continuous Offline Running

  • Optimize for Resource Usage: Stick with small models (tiny/base for Whisper, small models for Vosk) to keep CPU/memory usage low on edge devices.
  • Handle Audio Stream Stability: Use a robust audio library like PyAudio or sounddevice to avoid dropouts in continuous capture.
  • Add Error Handling: Include logic to ignore partial matches or low-confidence results to reduce false triggers.

内容的提问来源于stack exchange,提问作者arun

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:54:27