You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MoviePy VideoClip起止时间设置无效,字幕与语音不同步求助

问题描述

我想给视频添加精准字幕,让每个单词在其发音的准确时刻显示,操作步骤如下:

  1. 使用Whisper提取单词时间戳:
def get_words_per_time(audio_speech_file):
    model = whisper.load_model("base")
    transcribe = model.transcribe(
        audio=audio_speech_file, fp16=False, word_timestamps=True
    )
    segments = transcribe["segments"]
    words = []

    for seg in segments:
        for word in seg["words"]:
            words.append(
                {
                    "word": word["word"],
                    "start": word["start"],
                    "end": word["end"],
                    "prob": round(word["probability"], 4),
                }
            )
    return words
  1. 使用MoviePy生成双单词字幕片段:
def generate_captions(
    words,
    font="Komika",
    fontsize=32,
    color="White",
    align="center",
    stroke_width=3,
    stroke_color="black",
):
    text_comp = []
    for i in track(range(0, len(words), 2), description="Creating captions..."):
        word1 = words[i]
        if i + 1 < len(words):
            word2 = words[i + 1]
        text_clip = TextClip(
            f"{word1['word']} {word2['word'] if i + 1 < len(words) else ''}",
            font=font,  # Change Font if not found
            fontsize=fontsize,
            color=color,
            align=align,
            method="caption",
            size=(660, None),
            stroke_width=stroke_width,
            stroke_color=stroke_color,
        )
        text_clip = text_clip.set_start(word1["start"])
        text_clip = text_clip.set_end(
            word2["end"] if i + 1 < len(words) else word1["end"]
        )
        text_comp.append(text_clip)
    return text_comp
  1. 将字幕合并到原视频:
vid_clip = CompositeVideoClip(
    [vid_clip, concatenate_videoclips(text_comp).set_position(("center", 860))]
)

最终输出视频中,字幕流动速度明显快于语音,仿佛时间戳设置未生效。提取到的单词时间戳数据如下:

[
    {
        "word": "This",
        "start": 0.0,
        "end": 0.22,
        "prob": 0.805
    },
    {
        "word": "is",
        "start": 0.22,
        "end": 0.42,
        "prob": 0.9991
    },
    {
        "word": "a",
        "start": 0.42,
        "end": 0.6,
        "prob": 0.999
    },
    {
        "word": "test,",
        "start": 0.6,
        "end": 1.04,
        "prob": 0.9939
    },
    {
        "word": "to",
        "start": 1.18,
        "end": 1.3,
        "prob": 0.9847
    },
    {
        "word": "show",
        "start": 1.3,
        "end": 1.54,
        "prob": 0.9971
    },
    {
        "word": "words",
        "start": 1.54,
        "end": 1.9,
        "prob": 0.995
    },
    {
        "word": "does",
        "start": 1.9,
        "end": 2.16,
        "prob": 0.997
    },
    {
        "word": "not",
        "start": 2.16,
        "end": 2.4,
        "prob": 0.9978
    },
    {
        "word": "appear.",
        "start": 2.4,
        "end": 2.82,
        "prob": 0.9984
    },
    {
        "word": "At",
        "start": 3.46,
        "end": 3.6,
        "prob": 0.9793
    },
    {
        "word": "their",
        "start": 3.6,
        "end": 3.8,
        "prob": 0.9984
    },
    {
        "word": "proper",
        "start": 3.8,
        "end": 4.22,
        "prob": 0.9976
    },
    {
        "word": "time.",
        "start": 4.22,
        "end": 4.72,
        "prob": 0.999
    },
    {
        "word": "Thanks",
        "start": 5.04,
        "end": 5.4,
        "prob": 0.9662
    },
    {
        "word": "for,",
        "start": 5.4,
        "end": 5.66,
        "prob": 0.9941
    },
    {
        "word": "watching.",
        "start": 5.94,
        "end": 6.36,
        "prob": 0.7701
    }
]

请问造成该问题的原因可能是什么?

问题原因分析
  1. concatenate_videoclips的误用
    concatenate_videoclips的作用是将视频片段按顺序首尾拼接,会完全忽略每个片段预先设置的start和end时间戳。比如原本第一个字幕片段该在0.0-1.04秒显示,第二个在1.18-1.9秒显示,但拼接后第二个会紧接第一个结束就播放,跳过了中间的间隔,直接导致字幕整体播放速度加快。

正确做法是直接将所有字幕片段放入CompositeVideoClip,它会尊重每个片段的时间戳设置:

vid_clip = CompositeVideoClip(
    [vid_clip] + [clip.set_position(("center", 860)) for clip in text_comp]
)
  1. 字幕拼接逻辑不符合需求
    你将相邻两个单词合并为一个字幕片段,并把片段的时间设为第一个单词的开始到第二个单词的结束,这会让两个单词同时显示整个时间段,而非各自在发音时刻单独出现。如果要实现每个单词精准时刻显示,应该为每个单词单独创建TextClip,或者调整逻辑让字幕在单词切换时更新。

  2. 潜在的JSON数据格式错误
    你提供的单词数据中有两处语法错误:"word": 'test,和"word": 'for,存在引号不匹配和多余换行的问题,虽然可能是粘贴失误,但如果代码实际处理的是错误数据,也可能导致字幕内容或时间解析异常,建议检查数据正确性。


内容的提问来源于stack exchange,提问作者ernesto casco velazquez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:10:57