如何在Python中获取带时间戳的OpenAI JSON格式转录结果?
实现带时间戳的Whisper转录结果
要获取包含文本、起始时间和时长的结构化转录结果,只需修改OpenAI API调用参数并简单处理返回数据即可:
1. 修改API调用参数
在openai.Audio.transcribe方法中添加response_format="verbose_json"参数,该参数会让API返回包含分段时间信息的详细JSON数据。
2. 整理返回结果
API返回的详细数据里的segments字段,每个元素都包含text、start和duration三个所需字段,提取这些字段并重组为目标格式即可。
修改后的完整代码
import openai import json openai.organization = "org-xxxxxx" openai.api_key = "sk-xxxxx" audio_file_path = "/Users/tejaksha/Downloads/dhoni.mp4" # 调用API时指定返回详细JSON格式 audio_file = open(audio_file_path, "rb") transcript = openai.Audio.transcribe("whisper-1", audio_file, response_format="verbose_json") # 整理成目标结构化格式 structured_result = { "transcript": [ { "text": seg["text"].strip(), "start": seg["start"], "duration": seg["duration"] } for seg in transcript["segments"] ] } # 打印格式化后的结果 print(json.dumps(structured_result, indent=2))
输出示例
运行代码后会得到符合需求的结构化结果:
{ "transcript": [ { "text": "Flat back, just got a little tight to him, he was wagging for it, set up for the slower ball and punished it.", "start": 0.0, "duration": 7.12 }, { "text": "The one's going straight down the ground.", "start": 7.12, "duration": 2.88 }, { "text": "And MS Daini just taking control.", "start": 10.0, "duration": 3.2 } ] }
注意:需确保你的OpenAI Python版本在v0.27.0及以上,该版本开始支持verbose_json响应格式。
内容的提问来源于stack exchange,提问作者CloudExplorer
相关产品推荐
相关产品推荐

