You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否通过Whisper API生成.SRT格式的转录文件?

如何用Whisper API生成SRT格式转录文件

Whisper API本身不支持直接输出SRT格式,但可以通过获取带时间戳的详细转录数据,自行转换生成SRT文件,具体步骤如下:

  • 第一步:获取带时间戳的转录数据
    默认调用openai.Audio.transcribe只会返回纯文本,需要指定response_format="verbose_json"参数,获取包含每个语音分段的开始/结束时间、文本等详细信息。修改后的API调用代码示例:

    import openai
    
    openai.api_key = "你的API密钥"
    audio_file = open("audio.mp3", "rb")
    transcript = openai.Audio.transcribe(
        model="whisper-1",
        file=audio_file,
        response_format="verbose_json"
    )
    
  • 第二步:将转录数据转换为SRT格式
    SRT文件有固定格式:每段包含序号、时间范围、文本内容。编写逻辑提取transcript["segments"]中的数据,转换为符合规范的SRT内容:

    def format_srt_time(seconds):
        hours = int(seconds // 3600)
        minutes = int((seconds % 3600) // 60)
        secs = int(seconds % 60)
        millis = int((seconds - int(seconds)) * 1000)
        return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"
    
    def build_srt_content(transcript_data):
        srt_lines = []
        for i, segment in enumerate(transcript_data["segments"], 1):
            start = format_srt_time(segment["start"])
            end = format_srt_time(segment["end"])
            text = segment["text"].strip()
            srt_lines.extend([str(i), f"{start} --> {end}", text, ""])
        return "\n".join(srt_lines)
    
    # 生成并保存SRT文件
    srt_content = build_srt_content(transcript)
    with open("audio_transcript.srt", "w", encoding="utf-8") as f:
        f.write(srt_content)
    

这种方式不局限于Python,只要能调用Whisper API并解析JSON响应的语言都能实现。核心逻辑都是:请求详细格式的转录结果,提取分段时间和文本,按SRT规范拼接内容。

内容的提问来源于stack exchange,提问作者SazCoR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 05:32:37