You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正确高亮显示两个字符串的差异?

Python实现字符串差异的颜色高亮问题

需求说明

我想用Python代码高亮显示两个字符串之间的差异,具体示例如下:

示例1

原字符串:

sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates."
sentence2 = "I am enjoying the summer breeze on the beach while I am doing some pilates."

预期结果(星号标注部分需显示为红色):

I *am* enjoying the summer breeze on the beach while I *am doing* some pilates.

示例2

原字符串:

sentence1 = "My favourite season is Autumn while my sister's favourite season is Winter."
sentence2 = "My favourite season is Autumn, while my sister's favourite season is Winter."

预期结果(逗号为差异部分):

"My favourite season is Autumn*,* while my sister's favourite season is Winter." 

尝试的代码及问题

我写了以下代码:

sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates."
sentence2 = "I'm enjoying the summer breeze on the beach while I am doing some pilates."

# Split the sentences into words
words1 = sentence1.split()
words2 = sentence2.split()

# Find the index where the sentences differ
index_of_difference = next((i for i, (word1, word2) in enumerate(zip(words1, words2)) if word1 != word2), None)

# Highlight differing part "am doing" in red
highlighted_words = []
for i, (word1, word2) in enumerate(zip(words1, words2)):
    if i == index_of_difference:
        highlighted_words.append('\033[91m' + word2 + '\033[0m')
    else:
        highlighted_words.append(word2)

highlighted_sentence = ' '.join(highlighted_words)
print(highlighted_sentence)

但运行后得到的结果是:

I'm enjoying the summer breeze on the beach while I *am* doing some

和预期的不符,预期应该是:

I'm enjoying the summer breeze on the beach while I *am doing* some pilates.

解决方案

你的问题出在两个地方:一是仅处理了第一个差异单词,没识别到连续的差异片段;二是基于空格拆分单词的方式无法覆盖字符级的差异(比如示例2的逗号)。可以用Python内置的difflib库精准识别连续差异,再统一高亮:

完整代码

import difflib

def highlight_diff(source_str, target_str):
    # 初始化字符串匹配器
    matcher = difflib.SequenceMatcher(None, source_str, target_str)
    result_parts = []
    
    # 遍历所有操作码,处理不同类型的片段
    for tag, src_start, src_end, target_start, target_end in matcher.get_opcodes():
        if tag in ('replace', 'insert'):
            # 对差异片段添加ANSI红色高亮
            result_parts.append(f'\033[91m{target_str[target_start:target_end]}\033[0m')
        elif tag == 'equal':
            # 相同片段直接添加
            result_parts.append(target_str[target_start:target_end])
        # 忽略delete操作,因为我们以目标字符串为展示基准
    
    return ''.join(result_parts)

# 测试示例1
sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates."
sentence2 = "I am enjoying the summer breeze on the beach while I am doing some pilates."
print(highlight_diff(sentence1, sentence2))

# 测试示例2
sentence3 = "My favourite season is Autumn while my sister's favourite season is Winter."
sentence4 = "My favourite season is Autumn, while my sister's favourite season is Winter."
print(highlight_diff(sentence3, sentence4))

代码说明

  1. difflib.SequenceMatcher:专门用于对比两个序列(这里是字符串),能精准找出连续的相同或差异片段。
  2. get_opcodes():返回的每个元组包含操作类型和对应片段的索引,操作类型包括:
    • equal:两段字符串相同
    • replace:目标字符串替换了原字符串的片段
    • insert:目标字符串新增了片段
    • delete:原字符串有但目标字符串没有的片段(这里忽略,因为我们展示目标字符串)
  3. 高亮逻辑:对replace和insert类型的片段添加ANSI红色控制码,在终端中会显示为红色。
  4. 适配场景:既支持单词级的连续差异(示例1的am doing),也支持单个字符的差异(示例2的逗号)。

内容的提问来源于stack exchange,提问作者Oliver

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 18:46:13