You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解析Google Speech API返回的JSON格式语音识别结果?

Hey there! Let’s walk through how to parse that Google Speech-to-Text response you’ve got stored in your data variable. I’ll use Python for examples since it’s the go-to for this kind of JSON parsing task, but the core logic translates to other languages too.

Parsing Google Speech-to-Text LongRunningRecognize Response

First, a quick recap: your data variable is a Python dictionary (converted from the JSON response) from the LongRunningRecognize API—this is the endpoint used for longer audio files that require asynchronous processing.

Step 1: Verify the job completed

First, check if the recognition job finished successfully using the done flag:

if data['done']:
    print("Recognition job completed! Let's parse the results.")
else:
    print("Job is still in progress—check back later.")

Step 2: Extract the core response data

The actual transcription results are nested inside the response key. Let’s grab that first to simplify access:

recognition_response = data['response']

Step 3: Pull the transcript and confidence score

The results list holds all transcribed segments (for longer audio, there will be multiple entries here). For your example, there’s one result entry. Let’s extract the top transcript (the highest-confidence option, which is always first in the alternatives list):

# Get the first transcribed segment (adjust the index if you need later segments)
first_segment = recognition_response['results'][0]

# Extract the alternative transcript options
transcript_options = first_segment['alternatives']

# Grab the top transcript and its confidence score
top_transcript = transcript_options[0]['transcript']
confidence_score = transcript_options[0]['confidence']  # Your snippet cut this off, but this is the full key

print(f"Transcript: {top_transcript}")
print(f"Confidence Score: {confidence_score}")

Step 4: Access job metadata (optional)

If you need details about the job’s timeline or progress, use the metadata key, which follows the LongRunningRecognizeMetadata schema:

job_metadata = data['metadata']
progress = job_metadata['progressPercent']
start_time = job_metadata['startTime']
completion_time = job_metadata['lastUpdateTime']

print(f"Job started at: {start_time}")
print(f"Job finished at: {completion_time}")

Handling longer audio with multiple segments

For longer audio files, the results list will have multiple entries (each representing a distinct audio segment). You can loop through all segments to collect every transcript:

for segment_num, result in enumerate(recognition_response['results'], start=1):
    best_option = result['alternatives'][0]
    print(f"Segment {segment_num}: {best_option['transcript']} (Confidence: {best_option['confidence']})")

Pro Tips to Avoid Errors:

  • Use dict.get() to safely handle missing keys (e.g., if a transcript doesn’t have a confidence score):
    confidence_score = transcript_options[0].get('confidence', 'No score available')
    
  • The alternatives list includes all possible transcriptions sorted by confidence—you can access lower-confidence options if needed, but the first one is almost always the most accurate.

内容的提问来源于stack exchange,提问作者user9316498

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:39:20