You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python的difflib实现diff3?替代Windows下无法调用的diff3命令

用Python difflib实现跨平台三方文件合并(替代diff3)

由于Windows环境默认没有diff3命令,我们可以基于Python的difflib手动实现Git风格的三方合并逻辑——核心是对比BASE、HEAD、OTHER三个版本的差异,合并无冲突修改,标记冲突块。

替代实现代码

import difflib

def merge_blobs(o_base, o_HEAD, o_other):
    # 辅助函数:将对象内容转换为保留换行符的行列表
    def get_content_lines(oid):
        if not oid:
            return []
        raw_data = data.get_object(oid).decode("utf-8")
        # 保留换行符,避免拼接时丢失原格式
        return raw_data.splitlines(keepends=True)

    # 获取三个版本的行内容
    base_lines = get_content_lines(o_base)
    head_lines = get_content_lines(o_HEAD)
    other_lines = get_content_lines(o_other)

    # 初始化SequenceMatcher用于差异对比
    matcher_head = difflib.SequenceMatcher(None, base_lines, head_lines)
    matcher_other = difflib.SequenceMatcher(None, base_lines, other_lines)

    # 获取两个版本相对于BASE的修改操作序列(opcode)
    head_ops = list(matcher_head.get_opcodes())
    other_ops = list(matcher_other.get_opcodes())

    merged = []
    base_idx = head_idx = other_idx = 0

    while base_idx < len(base_lines):
        # 找到当前BASE位置对应的HEAD和OTHER修改操作
        current_head_op = next((op for op in head_ops if op[1] == base_idx), None)
        current_other_op = next((op for op in other_ops if op[1] == base_idx), None)

        # 情况1:两边都未修改当前BASE内容
        if not current_head_op and not current_other_op:
            merged.append(base_lines[base_idx])
            base_idx += 1
            head_idx += 1
            other_idx += 1
            continue

        # 情况2:仅HEAD修改了当前块
        if current_head_op and not current_other_op:
            tag, i1, i2, j1, j2 = current_head_op
            merged.extend(head_lines[j1:j2] if tag != 'equal' else base_lines[i1:i2])
            base_idx = i2
            head_idx = j2
            other_idx += (i2 - i1)
            continue

        # 情况3:仅OTHER修改了当前块
        if not current_head_op and current_other_op:
            tag, i1, i2, j1, j2 = current_other_op
            merged.extend(other_lines[j1:j2] if tag != 'equal' else base_lines[i1:i2])
            base_idx = i2
            head_idx += (i2 - i1)
            other_idx = j2
            continue

        # 情况4:两边都修改了当前块,检查冲突
        h_tag, h_i1, h_i2, h_j1, h_j2 = current_head_op
        o_tag, o_i1, o_i2, o_j1, o_j2 = current_other_op

        # 修改范围重叠,判定为冲突
        if not (h_i2 <= o_i1 or o_i2 <= h_i1):
            merged.append('<<<<<<< HEAD\n')
            merged.extend(head_lines[h_j1:h_j2])
            merged.append('=======\n')
            merged.extend(other_lines[o_j1:o_j2])
            merged.append('>>>>>>> MERGE_HEAD\n')
            # 跳过双方修改的块
            base_idx = max(h_i2, o_i2)
            head_idx = h_j2
            other_idx = o_j2
        else:
            # 修改不重叠,先处理位置靠前的修改
            if h_i1 < o_i1:
                merged.extend(head_lines[h_j1:h_j2] if h_tag != 'equal' else base_lines[h_i1:h_i2])
                base_idx = h_i2
                head_idx = h_j2
                other_idx += (h_i2 - h_i1)
            else:
                merged.extend(other_lines[o_j1:o_j2] if o_tag != 'equal' else base_lines[o_i1:o_i2])
                base_idx = o_i2
                head_idx += (o_i2 - o_i1)
                other_idx = o_j2

    # 处理HEAD或OTHER中超出BASE长度的内容
    merged.extend(head_lines[head_idx:])
    merged.extend(other_lines[other_idx:])

    # 合并为字符串并转字节返回
    return ''.join(merged).encode("utf-8")

代码说明

  1. 内容预处理:通过get_content_lines将对象内容转为保留换行符的行列表,避免拼接时破坏原文件格式。
  2. 差异对比:用difflib.SequenceMatcher生成BASE与HEAD、BASE与OTHER的修改操作序列(opcode),每个opcode包含修改类型(equal/replace/insert/delete)和位置信息。
  3. 合并逻辑:
    • 无修改:直接保留BASE内容
    • 单边修改:应用对应版本的修改
    • 双边修改:判断修改范围是否重叠,重叠则添加Git风格的冲突标记,不重叠则依次应用两边修改
  4. 收尾处理:合并HEAD或OTHER中超出BASE长度的新增内容。

与原diff3实现的对齐点

  • 生成的冲突块格式贴近Git风格,符合你的Git学习项目需求
  • 纯Python实现,跨平台支持Windows
  • 保留原代码的输入输出格式(接收oid,返回字节数据)

内容的提问来源于stack exchange,提问作者Doruk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 14:39:52