如何用Python版edge-tts生成带词级时间戳的字幕?
用edge-tts生成带单词级时间戳的字幕
edge-tts底层通过WebSocket与微软服务通信,返回的消息中包含单词级别的时间戳数据,无需依赖命令行工具,直接在Python中即可提取并生成字幕。
以下是同时生成音频和带单词时间戳SRT字幕的实现代码:
import edge_tts import asyncio import json from datetime import timedelta def format_time(nanoseconds: int) -> str: # 将微软返回的100纳秒单位时间,转换为SRT标准的时分秒.毫秒格式 td = timedelta(microseconds=nanoseconds // 10) hours = td.seconds // 3600 minutes = (td.seconds // 60) % 60 seconds = td.seconds % 60 milliseconds = td.microseconds // 1000 return f"{hours:02d}:{minutes:02d}:{seconds:02d},{milliseconds:03d}" async def generate_speech_with_subtitles(): text = "Hello, this is a test of the Edge TTS service." voice = "en-US-GuyNeural" audio_file = "output.mp3" subtitle_file = "output.srt" communicate = edge_tts.Communicate(text, voice) subtitles = [] current_index = 1 # 遍历WebSocket消息流,提取单词时间戳 async for message in communicate: if message["type"] == "Metadata": metadata = json.loads(message["data"]) for event in metadata["MetadataEvent"]: if event["Type"] == "WordBoundary": start_time = format_time(event["Data"]["AudioOffset"]) # 通过起始偏移量+时长计算结束时间 end_time = format_time(event["Data"]["AudioOffset"] + event["Data"]["Duration"]) word = event["Data"]["Text"] # 跳过空文本(如标点符号的边界事件) if word.strip(): subtitles.append({ "index": current_index, "start": start_time, "end": end_time, "text": word }) current_index += 1 # 保存音频文件 await communicate.save(audio_file) # 写入SRT字幕文件 with open(subtitle_file, "w", encoding="utf-8") as f: for sub in subtitles: f.write(f"{sub['index']}\n") f.write(f"{sub['start']} --> {sub['end']}\n") f.write(f"{sub['text']}\n\n") asyncio.run(generate_speech_with_subtitles())
核心要点:
format_time函数:适配微软返回的时间单位,转换为字幕标准格式- 筛选
Metadata类型消息:其中的WordBoundary事件包含每个单词的起始偏移、时长和文本内容 - 生成字幕:提取到的单词数据可直接写入SRT文件,也可根据需求调整为VTT等其他字幕格式
内容的提问来源于stack exchange,提问作者Hashir Nawaz
相关产品推荐
相关产品推荐

