如何用Python提取Google Cloud Speech-to-Text API的纯转录文本?
Python提取Google Cloud Speech-to-Text转录文本的简便方法
处理本地JSON文件的方案
如果已经把API返回的JSON保存到了本地文件,用Python的json模块就能快速提取需要的转录文本:
import json # 读取原始JSON文件 with open('speech_output.json', 'r', encoding='utf-8') as f: raw_data = json.load(f) # 提取所有转录结果 transcript_list = [] for result in raw_data['results']: # 遍历每个结果下的备选转录(如果只需要最准确的第一个,直接取alternatives[0]即可) for alt in result['alternatives']: transcript_list.append(alt['transcript']) # 打印所有转录文本,每行一条 print('\n'.join(transcript_list)) # 也可以保存为简化后的JSON文件 simplified_json = {'transcripts': transcript_list} with open('clean_transcripts.json', 'w', encoding='utf-8') as f: json.dump(simplified_json, f, indent=2)
直接处理API响应的方案
如果是用Google Cloud的Python SDK直接调用Speech-to-Text API,可以在获取响应后直接提取,不用先转成JSON:
from google.cloud import speech_v1p1beta1 as speech # 初始化客户端 client = speech.SpeechClient() # 这里假设你已经完成了音频识别请求,拿到了response对象 # 示例请求代码(可根据你的音频格式调整): # audio = speech.RecognitionAudio(uri="gs://your-bucket/audio-file.wav") # config = speech.RecognitionConfig( # encoding=speech.RecognitionConfig.AudioEncoding.LINEAR16, # sample_rate_hertz=16000, # language_code="en-US", # ) # response = client.recognize(config=config, audio=audio) # 提取最优转录结果 transcript_list = [] for result in response.results: # 取第一个备选结果(通常是置信度最高的) best_transcript = result.alternatives[0].transcript transcript_list.append(best_transcript) # 输出结果 print('\n'.join(transcript_list))
关键说明
- 如果不需要所有备选转录结果,只保留每个
result下的第一个alternative即可,这是API返回的置信度最高的转录文本 - 代码里的文件路径和参数(比如语言代码、音频编码)需要根据你的实际情况调整
内容的提问来源于stack exchange,提问作者Matthias
相关产品推荐
相关产品推荐

