You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python实现WhisperX转录文本与字幕脚本的对齐方案问询

解决WhisperX转录文本与原始脚本的精准对齐问题

问题背景

需要将WhisperX生成的带单词级时间戳的转录文本,映射到原始朗读脚本上以生成精准字幕,但两者存在以下不匹配情况:

  • 复合词拆分/合并:转录文本拆分复合词(如hive+mind对应脚本的hivemind),或脚本使用复合词(如carefree对应转录的care+free)
  • 拼写与标点差异:如转录的Brian,对应脚本的Bryan.,la la le loo,对应lalaleeloo
  • 大小写不一致:转录的diving对应脚本的Diving

现有基于SequenceMatcher或Jaro距离的逐词匹配方案未考虑上下文顺序约束和多对一/一对多的映射关系,容易出现错位匹配,效果不佳。

改进方案:带约束的动态规划序列对齐

核心思路是通过动态规划(DP)构建转录序列与脚本序列的最优对齐,同时允许多对一/一对多的映射,兼顾字符相似度和上下文顺序约束。

步骤1:预处理

先统一文本格式,减少无关差异:

  • 统一转为小写
  • 去除标点符号(或保留但降低其匹配权重)
  • 对连续重复音节(如la la le loo)合并为无空格形式,与脚本的lalaleeloo对齐

步骤2:动态规划对齐实现

定义dp[i][j]为转录前i个单词与脚本前j个单词的最小匹配成本,成本计算包含:

  • 字符相似度:使用Jaro-Winkler距离计算单个/多个转录单词与脚本单词的匹配度
  • 位置约束:限制对齐时的跳转范围(如不能跳过超过3个单词),避免跨上下文匹配
  • 映射类型:允许1对1、多对1、1对多的映射,对应不同的成本权重

以下是具体实现代码:

import numpy as np
from jellyfish import jaro_winkler_similarity

def preprocess(text):
    # 统一小写+去除标点
    return text.lower().translate(str.maketrans('', '', '.,!?'))

def calculate_match_score(trans_words, script_word):
    # 计算多个转录单词合并后与脚本单词的相似度
    combined_trans = ''.join([preprocess(w) for w in trans_words])
    processed_script = preprocess(script_word)
    return jaro_winkler_similarity(combined_trans, processed_script)

def align_transcription_to_script(transcription, script):
    n = len(transcription)
    m = len(script)
    
    # 初始化DP表,dp[i][j]存储最小成本和对齐路径
    dp = [[(float('inf'), None) for _ in range(m+1)] for __ in range(n+1)]
    dp[0][0] = (0, None)
    
    # 允许的最大跳转范围,避免跨太多单词匹配
    max_jump = 3
    
    for i in range(n+1):
        for j in range(m+1):
            if i == 0 and j == 0:
                continue
            current_cost, _ = dp[i][j]
            
            # 情况1:跳过当前转录单词
            if i > 0:
                cost = dp[i-1][j][0] + 1  # 跳过成本
                if cost < current_cost:
                    dp[i][j] = (cost, ('skip_trans', i-1, j))
            
            # 情况2:跳过当前脚本单词
            if j > 0:
                cost = dp[i][j-1][0] + 1  # 跳过成本
                if cost < current_cost:
                    dp[i][j] = (cost, ('skip_script', i, j-1))
            
            # 情况3:多对一映射(转录k个单词对应脚本1个单词)
            for k in range(1, min(i, max_jump)+1):
                trans_segment = transcription[i-k:i]
                score = calculate_match_score(trans_segment, script[j-1])
                # 相似度越高,成本越低
                cost = dp[i-k][j-1][0] + (1 - score)
                if cost < current_cost:
                    dp[i][j] = (cost, ('many_to_one', i-k, j-1, k))
    
    # 回溯获取对齐结果
    alignment = []
    i, j = n, m
    while i > 0 or j > 0:
        action = dp[i][j][1]
        if not action:
            break
        if action[0] == 'many_to_one':
            _, prev_i, prev_j, k = action
            # 标记转录的i-k到i-1个单词对应脚本的j-1个单词
            for idx in range(prev_i, i):
                alignment.append((idx, prev_j))
            i, j = prev_i, prev_j
        elif action[0] == 'skip_trans':
            _, prev_i, prev_j = action
            alignment.append((prev_i, None))  # 转录单词无对应脚本
            i, j = prev_i, prev_j
        elif action[0] == 'skip_script':
            _, prev_i, prev_j = action
            alignment.append((None, prev_j))  # 脚本单词无对应转录
            i, j = prev_i, prev_j
    
    # 反转对齐结果,得到从前往后的映射
    alignment.reverse()
    # 整理为转录单词到脚本单词的索引映射(匹配示例格式)
    trans_to_script = []
    for trans_idx, script_idx in alignment:
        if trans_idx is not None:
            trans_to_script.append(script_idx + 1 if script_idx is not None else None)
    
    return trans_to_script

# 示例测试
transcription = ['The', 'hive', 'mind', 'dips', 'and', 'dabbles,', 'care', 'free', 'diving', 'for', 'la', 'la', 'le', 'loo,', 'Brian,', 'bobbing', 'heads', 'and', 'bottoms', 'rolls,', 'webbed', 'feet', 'and', 'orange', 'beaks.']
script = ['The', 'hivemind', 'dips', 'and', 'dabbles,', 'carefree', 'Diving', 'for', 'lalaleeloo', 'Bryan.', 'Bobbing', 'heads', 'and', 'bottoms', 'rolls,', 'Webbed-feet', 'and', 'orange', 'beaks.']

indexes = align_transcription_to_script(transcription, script)
print(indexes)

步骤3:时间戳合并

根据对齐结果,将多个转录单词对应的脚本单词的时间戳合并:取第一个转录单词的start和最后一个转录单词的end作为脚本单词的时间范围。

def merge_timestamps(transcription_segments, alignment, script):
    script_timestamps = []
    current_script_idx = None
    current_start = None
    current_end = None
    
    # 整理转录单词的时间戳列表
    trans_words_with_ts = []
    for seg in transcription_segments:
        for word in seg['words']:
            trans_words_with_ts.append({
                'text': word['text'],
                'start': word['start'],
                'end': word['end']
            })
    
    for trans_idx, script_idx in enumerate(alignment):
        if script_idx is None:
            continue
        # 切换到新的脚本单词
        if script_idx != current_script_idx:
            if current_script_idx is not None:
                script_timestamps.append({
                    'text': script[current_script_idx],
                    'start': current_start,
                    'end': current_end
                })
            current_script_idx = script_idx
            current_start = trans_words_with_ts[trans_idx]['start']
            current_end = trans_words_with_ts[trans_idx]['end']
        else:
            # 更新结束时间为当前转录单词的结束时间
            current_end = trans_words_with_ts[trans_idx]['end']
    
    # 添加最后一个脚本单词
    if current_script_idx is not None:
        script_timestamps.append({
            'text': script[current_script_idx],
            'start': current_start,
            'end': current_end
        })
    
    return script_timestamps

关键优化点

  • 上下文约束:通过max_jump限制跳转范围,避免跨段落的错误匹配
  • 多对一映射支持:专门处理复合词拆分/合并、连续音节拆分的场景
  • 相似度加权:使用Jaro-Winkler距离(对短字符串匹配更友好)替代简单序列匹配,提升拼写差异场景的匹配精度

内容的提问来源于stack exchange,提问作者James Beanly

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 19:17:03