You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python调用Google Cloud Speech-to-Text:如何保存词与时间戳为JSON

回答

Hey there! Let's tackle your questions step by step:

通用问题:Python库中设置类似--format=json的参数

First off, the --format=json flag you know from the gcloud CLI is a command-line specific feature that handles formatting CLI output directly. The google.cloud.speech Python client library works differently: it returns native Python objects (protobuf-generated classes like SpeechRecognitionResult and WordInfo) instead of pre-formatted text or JSON.

Google doesn't offer a direct equivalent flag in the Python SDK because the goal is to give you full control over processing and structuring data once you receive the response. You're meant to extract the fields you need from the returned objects and format them as required—this flexibility lets you tailor output to your exact use case.

具体问题:生成带时间戳的字典格式JSON

Your "temporary hack" is totally valid, but we can refine it to create a more structured, usable JSON output that's easier to work with downstream. Instead of separate lists for words, start times, and end times, let's build a list of word objects (each containing the word and its timestamps), plus include the full transcript and confidence score for context.

Here's an optimized version of your code that outputs a clean, structured dictionary and saves it to a JSON file:

import io
import json
import argparse
from google.cloud import speech
from google.cloud.speech import enums
from google.cloud.speech import types

def transcribe_file_with_word_time_offsets(speech_file, language):
    """Transcribe the given audio file synchronously and output word time offsets as JSON."""
    print("Starting transcription...")
    client = speech.SpeechClient(credentials=credentials)

    with io.open(speech_file, 'rb') as audio_file:
        content = audio_file.read()

    audio = types.RecognitionAudio(content=content)
    config = types.RecognitionConfig(
        encoding=enums.RecognitionConfig.AudioEncoding.FLAC,
        language_code=language,
        enable_word_time_offsets=True
    )

    print("Recognizing audio...")
    response = client.recognize(config, audio)
    print("Recognition complete!")

    # Build structured output
    transcript_output = []
    for result in response.results:
        alternative = result.alternatives[0]
        entry = {
            "full_transcript": alternative.transcript,
            "confidence": alternative.confidence,
            "words": []
        }
        for word_info in alternative.words:
            entry["words"].append({
                "word": word_info.word,
                "start_time": word_info.start_time.seconds + word_info.start_time.nanos * 1e-9,
                "end_time": word_info.end_time.seconds + word_info.end_time.nanos * 1e-9
            })
        transcript_output.append(entry)

    # Save to JSON file
    with open("transcription_result.json", "w") as f:
        json.dump(transcript_output, f, indent=2)
    
    print("JSON output saved to transcription_result.json")

if __name__ == '__main__':
    parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
    parser.add_argument(dest='path', help='Audio file to be recognized')
    args = parser.parse_args()
    transcribe_file_with_word_time_offsets(args.path, 'en-US')

Key improvements:

  • Intuitive structure: Each result entry includes the full transcript, confidence score, and a list of word objects (each with word, start_time, end_time)
  • Direct JSON export: Uses Python's built-in json module to serialize the dictionary to a properly formatted, human-readable JSON file
  • Flexibility: You can easily add/remove fields from the output structure without modifying core API logic

If you ever need the full raw API response in JSON (including all fields returned by the Speech API), you can use the protobuf library's MessageToJson method:

from google.protobuf.json_format import MessageToJson

# After receiving the API response
raw_json_response = MessageToJson(response)
with open("raw_transcription_response.json", "w") as f:
    f.write(raw_json_response)

This is useful if you want to preserve every detail from the API, but the custom dictionary approach is cleaner for focused use cases like extracting word timestamps.


内容的提问来源于stack exchange,提问作者tmo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:41:28