You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Google Cloud Speech-to-Text API时无限等待问题求助

Troubleshooting Infinite Wait with Google Cloud Speech-to-Text API

Let's walk through the common issues that could be causing your long-running recognition call to hang, along with practical fixes for each:

1. Mismatched Audio Encoding Configuration

Your code specifies FLAC encoding, but you mentioned converting an MP3 to .raw—which is typically uncompressed PCM audio (not FLAC). This is a critical mismatch.

FLAC is a compressed, lossless format with its own header structure, while .raw files are raw PCM data with no header. For raw PCM audio, you need to use LINEAR16 encoding instead of FLAC.

2. Incorrect Sample Rate

Double-check that your converted .raw file actually uses a 16000 Hz sample rate. If your original MP3 was 44100 Hz or another rate and you didn't resample it during conversion, the API will struggle to process the audio, leading to hangs or silent failures.

You can verify the sample rate with ffmpeg:

ffmpeg -i gs://cloudh3-200314.appspot.com/cs.raw 2>&1 | grep "Sample rate"

3. Missing Cloud Storage Permissions

The Speech-to-Text service account needs explicit permission to read your audio file from Cloud Storage. If it doesn't have access, the operation will hang indefinitely waiting for file access.

To fix this:

  • Navigate to your Cloud Storage bucket's permissions settings
  • Add the service account service-<YOUR_PROJECT_NUMBER>@gcp-sa-speech.iam.gserviceaccount.com
  • Assign the Storage Object Viewer role to this account

4. Corrupted or Invalid Audio File

If the MP3-to-raw conversion failed (e.g., incomplete conversion, wrong parameters), your .raw file might be corrupted or empty. Validate the file with ffmpeg:

ffmpeg -i gs://cloudh3-200314.appspot.com/cs.raw -f null -

If you get errors here, re-convert your MP3 to raw with the correct parameters:

ffmpeg -i input.mp3 -acodec pcm_s16le -ar 16000 -ac 1 cs.raw

This command converts to 16-bit, 16000 Hz, mono raw audio—ideal for Speech-to-Text.

5. Add Timeout to Debug Hangs

Your current code uses operation.result() with no timeout, so it will wait forever. Adding a timeout will force an error to be thrown, giving you actionable debug info:

response = operation.result(timeout=300)  # Wait up to 5 minutes

Fixed Code Example

Here's your code updated with the correct encoding, timeout, and clarifying comments:

def transcribe_gcs(gcs_uri):
    """Asynchronously transcribes the audio file specified by the gcs_uri."""
    from google.cloud import speech
    from google.cloud.speech import enums
    from google.cloud.speech import types

    client = speech.SpeechClient()

    audio = types.RecognitionAudio(uri=gcs_uri)
    # Use LINEAR16 for raw PCM audio; match sample rate to your actual file
    config = types.RecognitionConfig(
        encoding=enums.RecognitionConfig.AudioEncoding.LINEAR16,
        sample_rate_hertz=16000,
        language_code='en-US'
    )

    operation = client.long_running_recognize(config, audio)
    print('Waiting for operation to complete... (timeout after 5 minutes)')
    # Add timeout to avoid infinite wait
    response = operation.result(timeout=300)

    for result in response.results:
        print(u'Transcript: {}'.format(result.alternatives[0].transcript))
        print('Confidence: {}'.format(result.alternatives[0].confidence))

transcribe_gcs("gs://cloudh3-200314.appspot.com/cs.raw")

Start by fixing the encoding mismatch—this is the most probable root cause. If that doesn't resolve the issue, check the sample rate, permissions, and file validity next.

内容的提问来源于stack exchange,提问作者Roland Iordache

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:01:46