You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用difflib对比多行字符串列表差异(未知增删改项)

多行字符串列表的差异对比解决方案

直接合并所有行对比会丢失多行字符串的块边界信息,导致diff错误关联不同块的行。核心解决思路是先识别两个列表中对应的多行字符串块(完全匹配、修改、新增、删除),再对每个对应块做行级差异对比。

实现步骤

  1. 优先匹配完全一致的多行字符串,避免误判修改
  2. 用字符串相似度匹配,识别被修改的对应块
  3. 标记纯新增和删除的块
  4. 对每个对应块执行行级diff

示例代码

import difflib

old_list = ["one\ntwo\nthree","four\nfive\nsix","seven\neight\nnine"]
new_list = ["four\nfifty\nsix","seven\neight\nnine","ten\neleven\ntwelve"]

# 1. 匹配完全一致的块
old_matched = set()
new_matched = set()
matches = []

for i, old_str in enumerate(old_list):
    for j, new_str in enumerate(new_list):
        if old_str == new_str and i not in old_matched and j not in new_matched:
            matches.append(("match", i, j, old_str, new_str))
            old_matched.add(i)
            new_matched.add(j)

# 2. 用相似度匹配识别修改的块
old_remaining = [(i, s) for i, s in enumerate(old_list) if i not in old_matched]
new_remaining = [(j, s) for j, s in enumerate(new_list) if j not in new_matched]

for old_idx, old_str in old_remaining:
    best_ratio = 0.0
    best_new_idx = None
    best_new_str = None
    for new_idx, new_str in new_remaining:
        if new_idx in new_matched:
            continue
        # 计算字符串相似度
        ratio = difflib.SequenceMatcher(None, old_str, new_str).ratio()
        if ratio > best_ratio and ratio > 0.5:  # 可根据需求调整相似度阈值
            best_ratio = ratio
            best_new_idx = new_idx
            best_new_str = new_str
    if best_new_idx is not None:
        matches.append(("similar", old_idx, best_new_idx, old_str, best_new_str))
        new_matched.add(best_new_idx)

# 3. 标记纯删除和新增的块
for old_idx, old_str in old_remaining:
    if old_idx not in {m[1] for m in matches}:
        matches.append(("deleted", old_idx, None, old_str, None))

for new_idx, new_str in new_remaining:
    if new_idx not in {m[2] for m in matches}:
        matches.append(("added", None, new_idx, None, new_str))

# 4. 生成每个块的行级差异
for match_type, old_idx, new_idx, old_str, new_str in matches:
    if match_type == "match":
        print(f"=== 完全匹配(原索引{old_idx} | 新索引{new_idx})===")
        print("无差异\n")
    elif match_type == "similar":
        print(f"=== 修改块(原索引{old_idx} | 新索引{new_idx})行级差异 ===")
        old_lines = old_str.splitlines()
        new_lines = new_str.splitlines()
        # 生成统一格式的diff
        diff = difflib.unified_diff(
            old_lines, new_lines,
            fromfile=f"old_block_{old_idx}",
            tofile=f"new_block_{new_idx}",
            lineterm=''
        )
        for line in diff:
            print(line)
        print()
    elif match_type == "deleted":
        print(f"=== 删除块(原索引{old_idx})===")
        print(old_str + "\n")
    elif match_type == "added":
        print(f"=== 新增块(新索引{new_idx})===")
        print(new_str + "\n")

代码说明

  • 完全匹配优先:确保未修改的块不会被误识别为修改
  • 相似度阈值:0.5可根据实际情况调整,值越高要求匹配度越强
  • 块级+行级diff:先定位块的变化类型,再深入行级差异,保留了原始结构信息

内容的提问来源于stack exchange,提问作者Moe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 01:50:30