Google Cloud语音转文字API超时:Jarvis助手长响应时崩溃
Let's break down why your app is crashing first: your main listening loop gets blocked when the assistant is processing and responding to a command. The original infinite stream logic relies on restarting the Google Speech-to-Text connection every 55 seconds (via STREAMING_LIMIT), but if your cmd.discover or response calls take longer than that, the loop can't trigger the restart in time—hitting Google's 60-second limit for single streaming sessions.
The Root Cause
In your current code, when listen_print_loop detects the wake word, it calls search() which runs your command handling (cmd.discover, cmd.respond) synchronously. This pauses the entire listening flow until the command finishes. If that process takes over 55 seconds, the stream never gets refreshed, leading to the API error and crash.
The Solution: Offload Command Handling to a Thread
We'll move the command processing logic to a separate thread so the main listening loop can keep running and refresh the stream on schedule. Here's the modified code with key changes:
#!/usr/bin/env python # Copyright 2018 Google LLC # Licensed under the Apache License, Version 2.0 (the "License"); # you may not use this file except in compliance with the License. # You may obtain a copy of the License at # http://www.apache.org/licenses/LICENSE-2.0 # Unless required by applicable law or agreed to in writing, software # distributed under the License is distributed on an "AS IS" BASIS, # WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. # See the License for the specific language governing permissions and # limitations under the License. """Google Cloud Speech API sample application using the streaming API.""" # [START speech_transcribe_infinite_streaming] from __future__ import division import time import re import sys import os import threading # Added for threading support from google.cloud import speech from pygame.mixer import * from googletrans import Translator translator = Translator() init() import pyaudio from six.moves import queue os.environ["GOOGLE_APPLICATION_CREDENTIALS"] = "C:\Users\mnauf\Desktop\rehandevice\key.json" from commands2 import commander cmd=commander() # Audio recording parameters STREAMING_LIMIT = 55000 # Keep this under 60k to avoid hitting API limits SAMPLE_RATE = 16000 CHUNK_SIZE = int(SAMPLE_RATE / 10) # 100ms def get_current_time(): return int(round(time.time() * 1000)) def duration_to_secs(duration): return duration.seconds + (duration.nanos / float(1e9)) class ResumableMicrophoneStream: """Opens a recording stream as a generator yielding the audio chunks.""" def __init__(self, rate, chunk_size): self._rate = rate self._chunk_size = chunk_size self._num_channels = 1 self._max_replay_secs = 5 # Create a thread-safe buffer of audio data self._buff = queue.Queue() self.closed = True self.start_time = get_current_time() self.is_processing_command = False # Add flag to track command state # 2 bytes in 16 bit samples self._bytes_per_sample = 2 * self._num_channels self._bytes_per_second = self._rate * self._bytes_per_sample self._bytes_per_chunk = (self._chunk_size * self._bytes_per_sample) self._chunks_per_second = ( self._bytes_per_second // self._bytes_per_chunk) def __enter__(self): self.closed = False self.is_processing_command = False # Reset flag when starting stream self._audio_interface = pyaudio.PyAudio() self._audio_stream = self._audio_interface.open( format=pyaudio.paInt16, channels=self._num_channels, rate=self._rate, input=True, frames_per_buffer=self._chunk_size, stream_callback=self._fill_buffer, ) return self def __exit__(self, type, value, traceback): self._audio_stream.stop_stream() self._audio_stream.close() self.closed = True self._buff.put(None) self._audio_interface.terminate() def _fill_buffer(self, in_data, *args, **kwargs): """Continuously collect data from the audio stream, into the buffer.""" if not self.is_processing_command: # Only buffer if not processing command self._buff.put(in_data) return None, pyaudio.paContinue def generator(self): while not self.closed: if get_current_time() - self.start_time > STREAMING_LIMIT: self.start_time = get_current_time() break chunk = self._buff.get() if chunk is None: return data = [chunk] while True: try: chunk = self._buff.get(block=False) if chunk is None: return data.append(chunk) except queue.Empty: break yield b''.join(data) def process_command(transcript, code, stream): """Handles command processing in a separate thread.""" stream.is_processing_command = True try: print("Your command: ", transcript) if "hindi assistant" in transcript.lower(): cmd.respond("Alright. Talk to me in urdu", code=code) main('ur-PK') elif "english assistant" in transcript.lower(): cmd.respond("Alright. Talk to me in English", code=code) main('en-US') cmd.discover(text=transcript, code=code) # Simulate long-running command (remove this in production) # time.sleep(70) finally: stream.is_processing_command = False # Reset flag when done def search(responses, stream, code): responses = (r for r in responses if ( r.results and r.results[0].alternatives)) num_chars_printed = 0 for response in responses: if not response.results: continue result = response.results[0] if not result.alternatives: continue top_alternative = result.alternatives[0] transcript = top_alternative.transcript if code == 'ur-PK': transcript = translator.translate(transcript).text overwrite_chars = ' ' * (num_chars_printed - len(transcript)) if not result.is_final: sys.stdout.write(transcript + overwrite_chars + '\r') sys.stdout.flush() num_chars_printed = len(transcript) else: # Start command processing in a separate thread threading.Thread(target=process_command, args=(transcript + overwrite_chars, code, stream)).start() break num_chars_printed = 0 def listen_print_loop(responses, stream, code): responses = (r for r in responses if ( r.results and r.results[0].alternatives)) music.load(r"C:\Users\mnauf\Desktop\rehandevice\coins.mp3") num_chars_printed = 0 for response in responses: if stream.is_processing_command: continue # Skip processing while handling a command if not response.results: continue result = response.results[0] if not result.alternatives: continue top_alternative = result.alternatives[0] transcript = top_alternative.transcript overwrite_chars = ' ' * (num_chars_printed - len(transcript)) if not result.is_final: sys.stdout.write(transcript + overwrite_chars + '\r') sys.stdout.flush() num_chars_printed = len(transcript) else: print("Listen print loop", transcript + overwrite_chars) if re.search(r'\b(hello)\b', transcript.lower(), re.I) or re.search(r'\b(ہیلو)\b', transcript, re.I): music.play() search(responses, stream, code) num_chars_printed = 0 def main(code): cmd.respond("I am Rayhaan dot A Eye. How can I help you?", code=code) client = speech.SpeechClient() config = speech.types.RecognitionConfig( encoding=speech.enums.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=SAMPLE_RATE, language_code=code, # Fixed: use the passed code instead of hardcoding 'en-US' max_alternatives=1, enable_word_time_offsets=True) streaming_config = speech.types.StreamingRecognitionConfig( config=config, interim_results=True) mic_manager = ResumableMicrophoneStream(SAMPLE_RATE, CHUNK_SIZE) print('Say "Quit" or "Exit" to terminate the program.') with mic_manager as stream: while not stream.closed: audio_generator = stream.generator() requests = (speech.types.StreamingRecognizeRequest( audio_content=content) for content in audio_generator) responses = client.streaming_recognize(streaming_config, requests) try: listen_print_loop(responses, stream, code) except Exception as e: print(f"Error occurred: {e}", file=sys.stderr) # Restart the stream instead of crashing continue if __name__ == '__main__': main('en-US') # [END speech_transcribe_infinite_streaming]
Key Changes Explained
- Threading Support: Added
threadingto run command processing in the background, so the main listening loop doesn't block. - Command State Flag: Added
is_processing_commandtoResumableMicrophoneStreamto pause audio buffering and recognition while handling a command (prevents accidental re-wake). - Fixed Language Code: In
main(), we now use the passedcodefor the recognition config instead of hardcoding 'en-US'—this fixes the language switching logic. - Improved Exception Handling: Replaced the broken
except:block with a proper catch that prints errors and restarts the stream. - Isolated Command Logic: Moved command handling to
process_command()which runs in a thread, ensuring the main loop can keep refreshing the speech stream every 55 seconds.
Additional Tips
- Keep
STREAMING_LIMITat 55000 (55 seconds) to leave a buffer before hitting Google's 60-second limit. - If your command responses involve long-running tasks (like playing a long audio file), make sure those tasks also run asynchronously to avoid blocking the command thread.
- Add a timeout or cancellation mechanism for commands if needed, to prevent orphaned threads.
内容的提问来源于stack exchange,提问作者Muhammad Naufil

