Python使用difflib对比文本时,如何处理无意义换行符问题
解决difflib对比文本时受无意义换行影响的误判问题
要解决PDF读取引入的无意义换行导致difflib误判的问题,核心是先预处理文本,合并拆分完整词汇的无效换行,同时保留段落/句子级的有效换行,再进行文本对比。
预处理逻辑
无意义换行通常是单词中途被拆分(比如arti\ncle),而有效换行是段落或句子的分隔。我们可以通过以下规则判断并合并:
- 如果当前行结尾不是空格(说明不是完整短语结尾)
- 且下一行开头不是大写字母、数字,也不是句子起始标点(如
。!?:)后的字符(说明不是新句子/段落开头) - 则将两行合并,去掉中间的换行符
修改后的代码实现
import difflib import re def clean_meaningless_newlines(text): lines = text.splitlines() cleaned = [] i = 0 while i < len(lines): current_line = lines[i].rstrip() # 去掉行尾多余空格,避免干扰判断 # 如果不是最后一行,判断是否需要合并下一行 if i < len(lines) - 1: next_line = lines[i+1].lstrip() # 去掉行首多余空格 # 判断条件:当前行结尾非空格,且下一行开头不是大写/数字/句子起始标点后的字符 current_end_not_space = len(current_line) > 0 and not current_line[-1].isspace() next_start_not_sentence = len(next_line) > 0 and ( not next_line[0].isupper() and not next_line[0].isdigit() and not re.match(r'^[.!?:]\s*.', next_line) ) if current_end_not_space and next_start_not_sentence: # 合并两行 cleaned.append(current_line + next_line) i += 2 continue # 不需要合并则直接添加当前行 cleaned.append(current_line) i += 1 # 把处理后的行用换行符连接,保留有效换行 return '\n'.join(cleaned) def compare_texts(old_text, new_text): # 先清理无意义换行 cleaned_old = clean_meaningless_newlines(old_text) cleaned_new = clean_meaningless_newlines(new_text) old_lines = cleaned_old.splitlines() new_lines = cleaned_new.splitlines() d = difflib.Differ() diff = d.compare(old_lines, new_lines) added_lines = [] deleted_lines = [] for line in diff: if line.startswith('+ '): added_lines.append(line[2:]) elif line.startswith('- '): deleted_lines.append(line[2:]) return added_lines, deleted_lines if __name__ == "__main__": old_text = "arti\ncle\nthis is some old text." new_text = "article\nthis is some new text." added_lines, deleted_lines = compare_texts(old_text, new_text) print("Added lines:") print('\n'.join(added_lines)) print("\nDeleted lines:") print('\n'.join(deleted_lines))
代码说明
clean_meaningless_newlines函数:遍历文本行,根据规则合并无意义换行,同时保留有效换行。- 对比前先对新旧文本做预处理,确保拆分单词的换行被合并,再用difflib对比处理后的文本。
- 测试示例中,
arti\ncle会被合并为article,因此不会被判定为差异,仅会检测到old和new的区别。
自定义调整
如果你的PDF有特殊换行规则(比如换行前带连字符arti-\ncle),可以修改合并逻辑,比如判断当前行结尾是连字符时,合并时去掉连字符:
# 在合并时添加连字符处理 if current_line.endswith('-'): cleaned.append(current_line[:-1] + next_line) else: cleaned.append(current_line + next_line)
内容的提问来源于stack exchange,提问作者John Smith
相关产品推荐
相关产品推荐

