如何用Python正确高亮显示两个字符串的差异?
Python实现字符串差异的颜色高亮问题
需求说明
我想用Python代码高亮显示两个字符串之间的差异,具体示例如下:
示例1
原字符串:
sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates." sentence2 = "I am enjoying the summer breeze on the beach while I am doing some pilates."
预期结果(星号标注部分需显示为红色):
I *am* enjoying the summer breeze on the beach while I *am doing* some pilates.
示例2
原字符串:
sentence1 = "My favourite season is Autumn while my sister's favourite season is Winter." sentence2 = "My favourite season is Autumn, while my sister's favourite season is Winter."
预期结果(逗号为差异部分):
"My favourite season is Autumn*,* while my sister's favourite season is Winter."
尝试的代码及问题
我写了以下代码:
sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates." sentence2 = "I'm enjoying the summer breeze on the beach while I am doing some pilates." # Split the sentences into words words1 = sentence1.split() words2 = sentence2.split() # Find the index where the sentences differ index_of_difference = next((i for i, (word1, word2) in enumerate(zip(words1, words2)) if word1 != word2), None) # Highlight differing part "am doing" in red highlighted_words = [] for i, (word1, word2) in enumerate(zip(words1, words2)): if i == index_of_difference: highlighted_words.append('\033[91m' + word2 + '\033[0m') else: highlighted_words.append(word2) highlighted_sentence = ' '.join(highlighted_words) print(highlighted_sentence)
但运行后得到的结果是:
I'm enjoying the summer breeze on the beach while I *am* doing some
和预期的不符,预期应该是:
I'm enjoying the summer breeze on the beach while I *am doing* some pilates.
解决方案
你的问题出在两个地方:一是仅处理了第一个差异单词,没识别到连续的差异片段;二是基于空格拆分单词的方式无法覆盖字符级的差异(比如示例2的逗号)。可以用Python内置的difflib库精准识别连续差异,再统一高亮:
完整代码
import difflib def highlight_diff(source_str, target_str): # 初始化字符串匹配器 matcher = difflib.SequenceMatcher(None, source_str, target_str) result_parts = [] # 遍历所有操作码,处理不同类型的片段 for tag, src_start, src_end, target_start, target_end in matcher.get_opcodes(): if tag in ('replace', 'insert'): # 对差异片段添加ANSI红色高亮 result_parts.append(f'\033[91m{target_str[target_start:target_end]}\033[0m') elif tag == 'equal': # 相同片段直接添加 result_parts.append(target_str[target_start:target_end]) # 忽略delete操作,因为我们以目标字符串为展示基准 return ''.join(result_parts) # 测试示例1 sentence1 = "I'm enjoying the summer breeze on the beach while I do some pilates." sentence2 = "I am enjoying the summer breeze on the beach while I am doing some pilates." print(highlight_diff(sentence1, sentence2)) # 测试示例2 sentence3 = "My favourite season is Autumn while my sister's favourite season is Winter." sentence4 = "My favourite season is Autumn, while my sister's favourite season is Winter." print(highlight_diff(sentence3, sentence4))
代码说明
difflib.SequenceMatcher:专门用于对比两个序列(这里是字符串),能精准找出连续的相同或差异片段。get_opcodes():返回的每个元组包含操作类型和对应片段的索引,操作类型包括:equal:两段字符串相同replace:目标字符串替换了原字符串的片段insert:目标字符串新增了片段delete:原字符串有但目标字符串没有的片段(这里忽略,因为我们展示目标字符串)
- 高亮逻辑:对
replace和insert类型的片段添加ANSI红色控制码,在终端中会显示为红色。 - 适配场景:既支持单词级的连续差异(示例1的
am doing),也支持单个字符的差异(示例2的逗号)。
内容的提问来源于stack exchange,提问作者Oliver
相关产品推荐
相关产品推荐

