You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将语音与文本的字符差异列表合并为连续差异对并生成对比DataFrame?

实现方案:合并差异片段并生成对比DataFrame

针对你的需求,这里提供两种高效的实现方案,核心是把分散的字符差异合并为连续片段,并结构化输出为DataFrame。

方案一:利用SequenceMatcher.get_opcodes()直接提取差异片段

difflib.SequenceMatcher自带的get_opcodes()方法会直接返回文本对比的操作指令(如替换、删除、插入、匹配),无需手动处理ndiff的零散结果,是最简洁的方案。

代码实现

import pandas as pd
from difflib import SequenceMatcher

speech = "chapter 1 it was a bright cold day in April and The clocks were striking 13 Winston Smith his chin nuzzled into his breast in an effort to escape the vile wind slipped quickly through the glass doors of victory Mansions though not quickly enough to prevent a swirl of gritty dust from entering along with him"
groundtruth =  "Chapter 1 It was a bright cold day in April, and the clocks were striking thirteen. Winston Smith, his chin nuzzled into his breast in an effort to escape the vile wind, slipped quickly through the glass doors of Victory Mansions, though not quickly enough to prevent a swirl of gritty dust from entering along with him."

# 归一化处理(和你原代码一致)
def normalize(text):
    return text.lower().replace('\n',' ').replace('.','').replace(',','').replace('-','').replace('_','')

speech_norm = normalize(speech)
groundtruth_norm = normalize(groundtruth)

# 初始化匹配器并获取操作码
matcher = SequenceMatcher(None, speech_norm, groundtruth_norm)
opcodes = matcher.get_opcodes()

# 提取差异片段并整理成列表
diff_records = []
for tag, i1, i2, j1, j2 in opcodes:
    if tag == 'equal':
        continue  # 跳过匹配的部分
    # 提取对应片段
    speech_segment = speech_norm[i1:i2]
    groundtruth_segment = groundtruth_norm[j1:j2]
    # 记录操作类型和片段
    diff_records.append({
        "操作类型": tag,
        "语音识别片段": speech_segment,
        "基准文本片段": groundtruth_segment
    })

# 生成DataFrame
diff_df = pd.DataFrame(diff_records)
print(diff_df)

说明

  • get_opcodes()返回的每个元组格式为(tag, i1, i2, j1, j2):
    • tag:操作类型,包括replace(替换)、delete(语音多了内容)、insert(语音少了内容)、equal(匹配)
    • i1,i2:语音文本中该操作对应的字符索引范围
    • j1,j2:基准文本中该操作对应的字符索引范围
  • 过滤掉equal类型后,直接提取对应片段即可得到连续的差异块,无需手动合并零散字符。

方案二:手动合并ndiff结果的零散差异

如果你需要基于已有的speech_diff列表处理,可以通过遍历分组的方式合并连续的同类型差异,再配对删除/插入片段。

代码实现

import pandas as pd

speech_diff = ['- 1', '- 3', '+ t', '+ h', '+ i', '+ r', '+ t', '+ e', '+ e', '+ n', '- w', '- o', '- l', '- e', '-  ', '- w', '- y', '-  ', '- s', '- m', '- e', '- t', '-  ', '- o', '- f', '+  ', '+ d', '+ u', '+ r']

# 分组连续的同类型差异
groups = []
current_group = []
current_type = None

for item in speech_diff:
    typ = item[0]
    char = item[2:]
    if typ != current_type:
        if current_group:
            groups.append((current_type, ''.join(current_group)))
        current_group = [char]
        current_type = typ
    else:
        current_group.append(char)
# 加入最后一组
if current_group:
    groups.append((current_type, ''.join(current_group)))

# 配对删除和插入片段(处理替换场景)
diff_records = []
i = 0
while i < len(groups):
    typ, seg = groups[i]
    if typ == '-' and i+1 < len(groups) and groups[i+1][0] == '+':
        # 替换场景:语音删除的片段对应基准插入的片段
        diff_records.append({
            "操作类型": "replace",
            "语音识别片段": seg,
            "基准文本片段": groups[i+1][1]
        })
        i += 2
    else:
        # 纯删除或纯插入
        diff_records.append({
            "操作类型": typ,
            "语音识别片段": seg if typ == '-' else "",
            "基准文本片段": seg if typ == '+' else ""
        })
        i += 1

# 生成DataFrame
diff_df = pd.DataFrame(diff_records)
print(diff_df)

说明

  1. 分组阶段:遍历speech_diff,把连续的-(语音删除)或+(基准插入)字符合并成完整片段。
  2. 配对阶段:将相邻的-和+片段配对为replace操作,单独的-或+标记为纯删除/插入。
  3. 最终生成的DataFrame会清晰展示每个差异块的对应关系。

内容的提问来源于stack exchange,提问作者salda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 21:36:25