Python调用Google Cloud Speech-to-Text:如何保存词与时间戳为JSON
Hey there! Let's tackle your questions step by step:
通用问题:Python库中设置类似--format=json的参数
First off, the --format=json flag you know from the gcloud CLI is a command-line specific feature that handles formatting CLI output directly. The google.cloud.speech Python client library works differently: it returns native Python objects (protobuf-generated classes like SpeechRecognitionResult and WordInfo) instead of pre-formatted text or JSON.
Google doesn't offer a direct equivalent flag in the Python SDK because the goal is to give you full control over processing and structuring data once you receive the response. You're meant to extract the fields you need from the returned objects and format them as required—this flexibility lets you tailor output to your exact use case.
具体问题:生成带时间戳的字典格式JSON
Your "temporary hack" is totally valid, but we can refine it to create a more structured, usable JSON output that's easier to work with downstream. Instead of separate lists for words, start times, and end times, let's build a list of word objects (each containing the word and its timestamps), plus include the full transcript and confidence score for context.
Here's an optimized version of your code that outputs a clean, structured dictionary and saves it to a JSON file:
import io import json import argparse from google.cloud import speech from google.cloud.speech import enums from google.cloud.speech import types def transcribe_file_with_word_time_offsets(speech_file, language): """Transcribe the given audio file synchronously and output word time offsets as JSON.""" print("Starting transcription...") client = speech.SpeechClient(credentials=credentials) with io.open(speech_file, 'rb') as audio_file: content = audio_file.read() audio = types.RecognitionAudio(content=content) config = types.RecognitionConfig( encoding=enums.RecognitionConfig.AudioEncoding.FLAC, language_code=language, enable_word_time_offsets=True ) print("Recognizing audio...") response = client.recognize(config, audio) print("Recognition complete!") # Build structured output transcript_output = [] for result in response.results: alternative = result.alternatives[0] entry = { "full_transcript": alternative.transcript, "confidence": alternative.confidence, "words": [] } for word_info in alternative.words: entry["words"].append({ "word": word_info.word, "start_time": word_info.start_time.seconds + word_info.start_time.nanos * 1e-9, "end_time": word_info.end_time.seconds + word_info.end_time.nanos * 1e-9 }) transcript_output.append(entry) # Save to JSON file with open("transcription_result.json", "w") as f: json.dump(transcript_output, f, indent=2) print("JSON output saved to transcription_result.json") if __name__ == '__main__': parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) parser.add_argument(dest='path', help='Audio file to be recognized') args = parser.parse_args() transcribe_file_with_word_time_offsets(args.path, 'en-US')
Key improvements:
- Intuitive structure: Each result entry includes the full transcript, confidence score, and a list of word objects (each with
word,start_time,end_time) - Direct JSON export: Uses Python's built-in
jsonmodule to serialize the dictionary to a properly formatted, human-readable JSON file - Flexibility: You can easily add/remove fields from the output structure without modifying core API logic
If you ever need the full raw API response in JSON (including all fields returned by the Speech API), you can use the protobuf library's MessageToJson method:
from google.protobuf.json_format import MessageToJson # After receiving the API response raw_json_response = MessageToJson(response) with open("raw_transcription_response.json", "w") as f: f.write(raw_json_response)
This is useful if you want to preserve every detail from the API, but the custom dictionary approach is cleaner for focused use cases like extracting word timestamps.
内容的提问来源于stack exchange,提问作者tmo

