You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python版edge-tts生成带词级时间戳的字幕?

用edge-tts生成带单词级时间戳的字幕

edge-tts底层通过WebSocket与微软服务通信,返回的消息中包含单词级别的时间戳数据,无需依赖命令行工具,直接在Python中即可提取并生成字幕。

以下是同时生成音频和带单词时间戳SRT字幕的实现代码:

import edge_tts
import asyncio
import json
from datetime import timedelta

def format_time(nanoseconds: int) -> str:
    # 将微软返回的100纳秒单位时间,转换为SRT标准的时分秒.毫秒格式
    td = timedelta(microseconds=nanoseconds // 10)
    hours = td.seconds // 3600
    minutes = (td.seconds // 60) % 60
    seconds = td.seconds % 60
    milliseconds = td.microseconds // 1000
    return f"{hours:02d}:{minutes:02d}:{seconds:02d},{milliseconds:03d}"

async def generate_speech_with_subtitles():
    text = "Hello, this is a test of the Edge TTS service."
    voice = "en-US-GuyNeural"
    audio_file = "output.mp3"
    subtitle_file = "output.srt"

    communicate = edge_tts.Communicate(text, voice)
    subtitles = []
    current_index = 1

    # 遍历WebSocket消息流,提取单词时间戳
    async for message in communicate:
        if message["type"] == "Metadata":
            metadata = json.loads(message["data"])
            for event in metadata["MetadataEvent"]:
                if event["Type"] == "WordBoundary":
                    start_time = format_time(event["Data"]["AudioOffset"])
                    # 通过起始偏移量+时长计算结束时间
                    end_time = format_time(event["Data"]["AudioOffset"] + event["Data"]["Duration"])
                    word = event["Data"]["Text"]
                    # 跳过空文本(如标点符号的边界事件)
                    if word.strip():
                        subtitles.append({
                            "index": current_index,
                            "start": start_time,
                            "end": end_time,
                            "text": word
                        })
                        current_index += 1

    # 保存音频文件
    await communicate.save(audio_file)

    # 写入SRT字幕文件
    with open(subtitle_file, "w", encoding="utf-8") as f:
        for sub in subtitles:
            f.write(f"{sub['index']}\n")
            f.write(f"{sub['start']} --> {sub['end']}\n")
            f.write(f"{sub['text']}\n\n")

asyncio.run(generate_speech_with_subtitles())

核心要点:

  • format_time函数:适配微软返回的时间单位,转换为字幕标准格式
  • 筛选Metadata类型消息:其中的WordBoundary事件包含每个单词的起始偏移、时长和文本内容
  • 生成字幕:提取到的单词数据可直接写入SRT文件,也可根据需求调整为VTT等其他字幕格式

内容的提问来源于stack exchange,提问作者Hashir Nawaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 17:15:01