MoviePy VideoClip起止时间设置无效,字幕与语音不同步求助
问题描述
我想给视频添加精准字幕,让每个单词在其发音的准确时刻显示,操作步骤如下:
- 使用Whisper提取单词时间戳:
def get_words_per_time(audio_speech_file): model = whisper.load_model("base") transcribe = model.transcribe( audio=audio_speech_file, fp16=False, word_timestamps=True ) segments = transcribe["segments"] words = [] for seg in segments: for word in seg["words"]: words.append( { "word": word["word"], "start": word["start"], "end": word["end"], "prob": round(word["probability"], 4), } ) return words
- 使用MoviePy生成双单词字幕片段:
def generate_captions( words, font="Komika", fontsize=32, color="White", align="center", stroke_width=3, stroke_color="black", ): text_comp = [] for i in track(range(0, len(words), 2), description="Creating captions..."): word1 = words[i] if i + 1 < len(words): word2 = words[i + 1] text_clip = TextClip( f"{word1['word']} {word2['word'] if i + 1 < len(words) else ''}", font=font, # Change Font if not found fontsize=fontsize, color=color, align=align, method="caption", size=(660, None), stroke_width=stroke_width, stroke_color=stroke_color, ) text_clip = text_clip.set_start(word1["start"]) text_clip = text_clip.set_end( word2["end"] if i + 1 < len(words) else word1["end"] ) text_comp.append(text_clip) return text_comp
- 将字幕合并到原视频:
vid_clip = CompositeVideoClip( [vid_clip, concatenate_videoclips(text_comp).set_position(("center", 860))] )
最终输出视频中,字幕流动速度明显快于语音,仿佛时间戳设置未生效。提取到的单词时间戳数据如下:
[ { "word": "This", "start": 0.0, "end": 0.22, "prob": 0.805 }, { "word": "is", "start": 0.22, "end": 0.42, "prob": 0.9991 }, { "word": "a", "start": 0.42, "end": 0.6, "prob": 0.999 }, { "word": "test,", "start": 0.6, "end": 1.04, "prob": 0.9939 }, { "word": "to", "start": 1.18, "end": 1.3, "prob": 0.9847 }, { "word": "show", "start": 1.3, "end": 1.54, "prob": 0.9971 }, { "word": "words", "start": 1.54, "end": 1.9, "prob": 0.995 }, { "word": "does", "start": 1.9, "end": 2.16, "prob": 0.997 }, { "word": "not", "start": 2.16, "end": 2.4, "prob": 0.9978 }, { "word": "appear.", "start": 2.4, "end": 2.82, "prob": 0.9984 }, { "word": "At", "start": 3.46, "end": 3.6, "prob": 0.9793 }, { "word": "their", "start": 3.6, "end": 3.8, "prob": 0.9984 }, { "word": "proper", "start": 3.8, "end": 4.22, "prob": 0.9976 }, { "word": "time.", "start": 4.22, "end": 4.72, "prob": 0.999 }, { "word": "Thanks", "start": 5.04, "end": 5.4, "prob": 0.9662 }, { "word": "for,", "start": 5.4, "end": 5.66, "prob": 0.9941 }, { "word": "watching.", "start": 5.94, "end": 6.36, "prob": 0.7701 } ]
请问造成该问题的原因可能是什么?
问题原因分析
concatenate_videoclips的误用concatenate_videoclips的作用是将视频片段按顺序首尾拼接,会完全忽略每个片段预先设置的start和end时间戳。比如原本第一个字幕片段该在0.0-1.04秒显示,第二个在1.18-1.9秒显示,但拼接后第二个会紧接第一个结束就播放,跳过了中间的间隔,直接导致字幕整体播放速度加快。
正确做法是直接将所有字幕片段放入CompositeVideoClip,它会尊重每个片段的时间戳设置:
vid_clip = CompositeVideoClip( [vid_clip] + [clip.set_position(("center", 860)) for clip in text_comp] )
字幕拼接逻辑不符合需求
你将相邻两个单词合并为一个字幕片段,并把片段的时间设为第一个单词的开始到第二个单词的结束,这会让两个单词同时显示整个时间段,而非各自在发音时刻单独出现。如果要实现每个单词精准时刻显示,应该为每个单词单独创建TextClip,或者调整逻辑让字幕在单词切换时更新。潜在的JSON数据格式错误
你提供的单词数据中有两处语法错误:"word": 'test,和"word": 'for,存在引号不匹配和多余换行的问题,虽然可能是粘贴失误,但如果代码实际处理的是错误数据,也可能导致字幕内容或时间解析异常,建议检查数据正确性。
内容的提问来源于stack exchange,提问作者ernesto casco velazquez
相关产品推荐
相关产品推荐

