使用Google Cloud Speech-to-Text API时无限等待问题求助
Let's walk through the common issues that could be causing your long-running recognition call to hang, along with practical fixes for each:
1. Mismatched Audio Encoding Configuration
Your code specifies FLAC encoding, but you mentioned converting an MP3 to .raw—which is typically uncompressed PCM audio (not FLAC). This is a critical mismatch.
FLAC is a compressed, lossless format with its own header structure, while .raw files are raw PCM data with no header. For raw PCM audio, you need to use LINEAR16 encoding instead of FLAC.
2. Incorrect Sample Rate
Double-check that your converted .raw file actually uses a 16000 Hz sample rate. If your original MP3 was 44100 Hz or another rate and you didn't resample it during conversion, the API will struggle to process the audio, leading to hangs or silent failures.
You can verify the sample rate with ffmpeg:
ffmpeg -i gs://cloudh3-200314.appspot.com/cs.raw 2>&1 | grep "Sample rate"
3. Missing Cloud Storage Permissions
The Speech-to-Text service account needs explicit permission to read your audio file from Cloud Storage. If it doesn't have access, the operation will hang indefinitely waiting for file access.
To fix this:
- Navigate to your Cloud Storage bucket's permissions settings
- Add the service account
service-<YOUR_PROJECT_NUMBER>@gcp-sa-speech.iam.gserviceaccount.com - Assign the
Storage Object Viewerrole to this account
4. Corrupted or Invalid Audio File
If the MP3-to-raw conversion failed (e.g., incomplete conversion, wrong parameters), your .raw file might be corrupted or empty. Validate the file with ffmpeg:
ffmpeg -i gs://cloudh3-200314.appspot.com/cs.raw -f null -
If you get errors here, re-convert your MP3 to raw with the correct parameters:
ffmpeg -i input.mp3 -acodec pcm_s16le -ar 16000 -ac 1 cs.raw
This command converts to 16-bit, 16000 Hz, mono raw audio—ideal for Speech-to-Text.
5. Add Timeout to Debug Hangs
Your current code uses operation.result() with no timeout, so it will wait forever. Adding a timeout will force an error to be thrown, giving you actionable debug info:
response = operation.result(timeout=300) # Wait up to 5 minutes
Fixed Code Example
Here's your code updated with the correct encoding, timeout, and clarifying comments:
def transcribe_gcs(gcs_uri): """Asynchronously transcribes the audio file specified by the gcs_uri.""" from google.cloud import speech from google.cloud.speech import enums from google.cloud.speech import types client = speech.SpeechClient() audio = types.RecognitionAudio(uri=gcs_uri) # Use LINEAR16 for raw PCM audio; match sample rate to your actual file config = types.RecognitionConfig( encoding=enums.RecognitionConfig.AudioEncoding.LINEAR16, sample_rate_hertz=16000, language_code='en-US' ) operation = client.long_running_recognize(config, audio) print('Waiting for operation to complete... (timeout after 5 minutes)') # Add timeout to avoid infinite wait response = operation.result(timeout=300) for result in response.results: print(u'Transcript: {}'.format(result.alternatives[0].transcript)) print('Confidence: {}'.format(result.alternatives[0].confidence)) transcribe_gcs("gs://cloudh3-200314.appspot.com/cs.raw")
Start by fixing the encoding mismatch—this is the most probable root cause. If that doesn't resolve the issue, check the sample rate, permissions, and file validity next.
内容的提问来源于stack exchange,提问作者Roland Iordache

