基于Python实现WhisperX转录文本与字幕脚本的对齐方案问询
解决WhisperX转录文本与原始脚本的精准对齐问题
问题背景
需要将WhisperX生成的带单词级时间戳的转录文本,映射到原始朗读脚本上以生成精准字幕,但两者存在以下不匹配情况:
- 复合词拆分/合并:转录文本拆分复合词(如
hive+mind对应脚本的hivemind),或脚本使用复合词(如carefree对应转录的care+free) - 拼写与标点差异:如转录的
Brian,对应脚本的Bryan.,la la le loo,对应lalaleeloo - 大小写不一致:转录的
diving对应脚本的Diving
现有基于SequenceMatcher或Jaro距离的逐词匹配方案未考虑上下文顺序约束和多对一/一对多的映射关系,容易出现错位匹配,效果不佳。
改进方案:带约束的动态规划序列对齐
核心思路是通过动态规划(DP)构建转录序列与脚本序列的最优对齐,同时允许多对一/一对多的映射,兼顾字符相似度和上下文顺序约束。
步骤1:预处理
先统一文本格式,减少无关差异:
- 统一转为小写
- 去除标点符号(或保留但降低其匹配权重)
- 对连续重复音节(如
la la le loo)合并为无空格形式,与脚本的lalaleeloo对齐
步骤2:动态规划对齐实现
定义dp[i][j]为转录前i个单词与脚本前j个单词的最小匹配成本,成本计算包含:
- 字符相似度:使用Jaro-Winkler距离计算单个/多个转录单词与脚本单词的匹配度
- 位置约束:限制对齐时的跳转范围(如不能跳过超过3个单词),避免跨上下文匹配
- 映射类型:允许1对1、多对1、1对多的映射,对应不同的成本权重
以下是具体实现代码:
import numpy as np from jellyfish import jaro_winkler_similarity def preprocess(text): # 统一小写+去除标点 return text.lower().translate(str.maketrans('', '', '.,!?')) def calculate_match_score(trans_words, script_word): # 计算多个转录单词合并后与脚本单词的相似度 combined_trans = ''.join([preprocess(w) for w in trans_words]) processed_script = preprocess(script_word) return jaro_winkler_similarity(combined_trans, processed_script) def align_transcription_to_script(transcription, script): n = len(transcription) m = len(script) # 初始化DP表,dp[i][j]存储最小成本和对齐路径 dp = [[(float('inf'), None) for _ in range(m+1)] for __ in range(n+1)] dp[0][0] = (0, None) # 允许的最大跳转范围,避免跨太多单词匹配 max_jump = 3 for i in range(n+1): for j in range(m+1): if i == 0 and j == 0: continue current_cost, _ = dp[i][j] # 情况1:跳过当前转录单词 if i > 0: cost = dp[i-1][j][0] + 1 # 跳过成本 if cost < current_cost: dp[i][j] = (cost, ('skip_trans', i-1, j)) # 情况2:跳过当前脚本单词 if j > 0: cost = dp[i][j-1][0] + 1 # 跳过成本 if cost < current_cost: dp[i][j] = (cost, ('skip_script', i, j-1)) # 情况3:多对一映射(转录k个单词对应脚本1个单词) for k in range(1, min(i, max_jump)+1): trans_segment = transcription[i-k:i] score = calculate_match_score(trans_segment, script[j-1]) # 相似度越高,成本越低 cost = dp[i-k][j-1][0] + (1 - score) if cost < current_cost: dp[i][j] = (cost, ('many_to_one', i-k, j-1, k)) # 回溯获取对齐结果 alignment = [] i, j = n, m while i > 0 or j > 0: action = dp[i][j][1] if not action: break if action[0] == 'many_to_one': _, prev_i, prev_j, k = action # 标记转录的i-k到i-1个单词对应脚本的j-1个单词 for idx in range(prev_i, i): alignment.append((idx, prev_j)) i, j = prev_i, prev_j elif action[0] == 'skip_trans': _, prev_i, prev_j = action alignment.append((prev_i, None)) # 转录单词无对应脚本 i, j = prev_i, prev_j elif action[0] == 'skip_script': _, prev_i, prev_j = action alignment.append((None, prev_j)) # 脚本单词无对应转录 i, j = prev_i, prev_j # 反转对齐结果,得到从前往后的映射 alignment.reverse() # 整理为转录单词到脚本单词的索引映射(匹配示例格式) trans_to_script = [] for trans_idx, script_idx in alignment: if trans_idx is not None: trans_to_script.append(script_idx + 1 if script_idx is not None else None) return trans_to_script # 示例测试 transcription = ['The', 'hive', 'mind', 'dips', 'and', 'dabbles,', 'care', 'free', 'diving', 'for', 'la', 'la', 'le', 'loo,', 'Brian,', 'bobbing', 'heads', 'and', 'bottoms', 'rolls,', 'webbed', 'feet', 'and', 'orange', 'beaks.'] script = ['The', 'hivemind', 'dips', 'and', 'dabbles,', 'carefree', 'Diving', 'for', 'lalaleeloo', 'Bryan.', 'Bobbing', 'heads', 'and', 'bottoms', 'rolls,', 'Webbed-feet', 'and', 'orange', 'beaks.'] indexes = align_transcription_to_script(transcription, script) print(indexes)
步骤3:时间戳合并
根据对齐结果,将多个转录单词对应的脚本单词的时间戳合并:取第一个转录单词的start和最后一个转录单词的end作为脚本单词的时间范围。
def merge_timestamps(transcription_segments, alignment, script): script_timestamps = [] current_script_idx = None current_start = None current_end = None # 整理转录单词的时间戳列表 trans_words_with_ts = [] for seg in transcription_segments: for word in seg['words']: trans_words_with_ts.append({ 'text': word['text'], 'start': word['start'], 'end': word['end'] }) for trans_idx, script_idx in enumerate(alignment): if script_idx is None: continue # 切换到新的脚本单词 if script_idx != current_script_idx: if current_script_idx is not None: script_timestamps.append({ 'text': script[current_script_idx], 'start': current_start, 'end': current_end }) current_script_idx = script_idx current_start = trans_words_with_ts[trans_idx]['start'] current_end = trans_words_with_ts[trans_idx]['end'] else: # 更新结束时间为当前转录单词的结束时间 current_end = trans_words_with_ts[trans_idx]['end'] # 添加最后一个脚本单词 if current_script_idx is not None: script_timestamps.append({ 'text': script[current_script_idx], 'start': current_start, 'end': current_end }) return script_timestamps
关键优化点
- 上下文约束:通过
max_jump限制跳转范围,避免跨段落的错误匹配 - 多对一映射支持:专门处理复合词拆分/合并、连续音节拆分的场景
- 相似度加权:使用Jaro-Winkler距离(对短字符串匹配更友好)替代简单序列匹配,提升拼写差异场景的匹配精度
内容的提问来源于stack exchange,提问作者James Beanly
相关产品推荐
相关产品推荐

