You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Cloud语音转文字API超时:Jarvis助手长响应时崩溃

Fix for Jarvis-like Voice Assistant Crashing When Response Exceeds 60 Seconds

Let's break down why your app is crashing first: your main listening loop gets blocked when the assistant is processing and responding to a command. The original infinite stream logic relies on restarting the Google Speech-to-Text connection every 55 seconds (via STREAMING_LIMIT), but if your cmd.discover or response calls take longer than that, the loop can't trigger the restart in time—hitting Google's 60-second limit for single streaming sessions.

The Root Cause

In your current code, when listen_print_loop detects the wake word, it calls search() which runs your command handling (cmd.discover, cmd.respond) synchronously. This pauses the entire listening flow until the command finishes. If that process takes over 55 seconds, the stream never gets refreshed, leading to the API error and crash.

The Solution: Offload Command Handling to a Thread

We'll move the command processing logic to a separate thread so the main listening loop can keep running and refresh the stream on schedule. Here's the modified code with key changes:

#!/usr/bin/env python
# Copyright 2018 Google LLC
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#     http://www.apache.org/licenses/LICENSE-2.0
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
"""Google Cloud Speech API sample application using the streaming API."""
# [START speech_transcribe_infinite_streaming]
from __future__ import division

import time
import re
import sys
import os
import threading  # Added for threading support

from google.cloud import speech
from pygame.mixer import *
from googletrans import Translator
translator = Translator()
init()
import pyaudio
from six.moves import queue

os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = "C:\Users\mnauf\Desktop\rehandevice\key.json"

from commands2 import commander
cmd=commander()

# Audio recording parameters
STREAMING_LIMIT = 55000  # Keep this under 60k to avoid hitting API limits
SAMPLE_RATE = 16000
CHUNK_SIZE = int(SAMPLE_RATE / 10)  # 100ms

def get_current_time():
    return int(round(time.time() * 1000))

def duration_to_secs(duration):
    return duration.seconds + (duration.nanos / float(1e9))

class ResumableMicrophoneStream:
    """Opens a recording stream as a generator yielding the audio chunks."""
    def __init__(self, rate, chunk_size):
        self._rate = rate
        self._chunk_size = chunk_size
        self._num_channels = 1
        self._max_replay_secs = 5

        # Create a thread-safe buffer of audio data
        self._buff = queue.Queue()
        self.closed = True
        self.start_time = get_current_time()
        self.is_processing_command = False  # Add flag to track command state

        # 2 bytes in 16 bit samples
        self._bytes_per_sample = 2 * self._num_channels
        self._bytes_per_second = self._rate * self._bytes_per_sample
        self._bytes_per_chunk = (self._chunk_size * self._bytes_per_sample)
        self._chunks_per_second = (
            self._bytes_per_second // self._bytes_per_chunk)

    def __enter__(self):
        self.closed = False
        self.is_processing_command = False  # Reset flag when starting stream

        self._audio_interface = pyaudio.PyAudio()
        self._audio_stream = self._audio_interface.open(
            format=pyaudio.paInt16,
            channels=self._num_channels,
            rate=self._rate,
            input=True,
            frames_per_buffer=self._chunk_size,
            stream_callback=self._fill_buffer,
        )

        return self

    def __exit__(self, type, value, traceback):
        self._audio_stream.stop_stream()
        self._audio_stream.close()
        self.closed = True
        self._buff.put(None)
        self._audio_interface.terminate()

    def _fill_buffer(self, in_data, *args, **kwargs):
        """Continuously collect data from the audio stream, into the buffer."""
        if not self.is_processing_command:  # Only buffer if not processing command
            self._buff.put(in_data)
        return None, pyaudio.paContinue

    def generator(self):
        while not self.closed:
            if get_current_time() - self.start_time > STREAMING_LIMIT:
                self.start_time = get_current_time()
                break
            chunk = self._buff.get()
            if chunk is None:
                return
            data = [chunk]

            while True:
                try:
                    chunk = self._buff.get(block=False)
                    if chunk is None:
                        return
                    data.append(chunk)
                except queue.Empty:
                    break
            yield b''.join(data)

def process_command(transcript, code, stream):
    """Handles command processing in a separate thread."""
    stream.is_processing_command = True
    try:
        print("Your command: ", transcript)
        if "hindi assistant" in transcript.lower():
            cmd.respond("Alright. Talk to me in urdu", code=code)
            main('ur-PK')
        elif "english assistant" in transcript.lower():
            cmd.respond("Alright. Talk to me in English", code=code)
            main('en-US')
        cmd.discover(text=transcript, code=code)
        # Simulate long-running command (remove this in production)
        # time.sleep(70)
    finally:
        stream.is_processing_command = False  # Reset flag when done

def search(responses, stream, code):
    responses = (r for r in responses if (
        r.results and r.results[0].alternatives))

    num_chars_printed = 0
    for response in responses:
        if not response.results:
            continue

        result = response.results[0]
        if not result.alternatives:
            continue

        top_alternative = result.alternatives[0]
        transcript = top_alternative.transcript

        if code == 'ur-PK':
            transcript = translator.translate(transcript).text

        overwrite_chars = ' ' * (num_chars_printed - len(transcript))

        if not result.is_final:
            sys.stdout.write(transcript + overwrite_chars + '\r')
            sys.stdout.flush()
            num_chars_printed = len(transcript)
        else:
            # Start command processing in a separate thread
            threading.Thread(target=process_command, args=(transcript + overwrite_chars, code, stream)).start()
            break
            num_chars_printed = 0

def listen_print_loop(responses, stream, code):
    responses = (r for r in responses if (
        r.results and r.results[0].alternatives))
    music.load(r"C:\Users\mnauf\Desktop\rehandevice\coins.mp3")
    num_chars_printed = 0
    for response in responses:
        if stream.is_processing_command:
            continue  # Skip processing while handling a command

        if not response.results:
            continue

        result = response.results[0]
        if not result.alternatives:
            continue

        top_alternative = result.alternatives[0]
        transcript = top_alternative.transcript

        overwrite_chars = ' ' * (num_chars_printed - len(transcript))

        if not result.is_final:
            sys.stdout.write(transcript + overwrite_chars + '\r')
            sys.stdout.flush()
            num_chars_printed = len(transcript)
        else:
            print("Listen print loop", transcript + overwrite_chars)
            if re.search(r'\b(hello)\b', transcript.lower(), re.I) or re.search(r'\b(ہیلو)\b', transcript, re.I):
                music.play()
                search(responses, stream, code)
            num_chars_printed = 0

def main(code):
    cmd.respond("I am Rayhaan dot A Eye. How can I help you?", code=code)
    client = speech.SpeechClient()
    config = speech.types.RecognitionConfig(
        encoding=speech.enums.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=SAMPLE_RATE,
        language_code=code,  # Fixed: use the passed code instead of hardcoding 'en-US'
        max_alternatives=1,
        enable_word_time_offsets=True)
    streaming_config = speech.types.StreamingRecognitionConfig(
        config=config,
        interim_results=True)

    mic_manager = ResumableMicrophoneStream(SAMPLE_RATE, CHUNK_SIZE)
    print('Say "Quit" or "Exit" to terminate the program.')

    with mic_manager as stream:
        while not stream.closed:
            audio_generator = stream.generator()
            requests = (speech.types.StreamingRecognizeRequest(
                audio_content=content)
                for content in audio_generator)

            responses = client.streaming_recognize(streaming_config, requests)

            try:
                listen_print_loop(responses, stream, code)
            except Exception as e:
                print(f"Error occurred: {e}", file=sys.stderr)
                # Restart the stream instead of crashing
                continue

if __name__ == '__main__':
    main('en-US')
# [END speech_transcribe_infinite_streaming]

Key Changes Explained

  1. Threading Support: Added threading to run command processing in the background, so the main listening loop doesn't block.
  2. Command State Flag: Added is_processing_command to ResumableMicrophoneStream to pause audio buffering and recognition while handling a command (prevents accidental re-wake).
  3. Fixed Language Code: In main(), we now use the passed code for the recognition config instead of hardcoding 'en-US'—this fixes the language switching logic.
  4. Improved Exception Handling: Replaced the broken except: block with a proper catch that prints errors and restarts the stream.
  5. Isolated Command Logic: Moved command handling to process_command() which runs in a thread, ensuring the main loop can keep refreshing the speech stream every 55 seconds.

Additional Tips

  • Keep STREAMING_LIMIT at 55000 (55 seconds) to leave a buffer before hitting Google's 60-second limit.
  • If your command responses involve long-running tasks (like playing a long audio file), make sure those tasks also run asynchronously to avoid blocking the command thread.
  • Add a timeout or cancellation mechanism for commands if needed, to prevent orphaned threads.

内容的提问来源于stack exchange,提问作者Muhammad Naufil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:06:19