使用set与strip()处理文本文件去重失效,请求排查修正代码
文本去重程序重复行问题排查与修正
问题场景
需求:使用Python上下文管理器读取input.txt,忽略首尾空白判定重复行,去重后写入output.txt并处理异常。
示例input.txt内容:
Hello, World!
Python is fun.
Python is fun.
Python is fun.
期望output.txt结果:
Hello, World!
Python is fun.
但运行代码后,实际output.txt出现重复行:
Hello, World!
Python is fun.
Python is fun.
用户提供的代码:
def remove_duplicates(input_file, output_file): try: with open(input_file, 'r') as in_file: lines = in_file.readlines() unique_lines = set(line.strip() for line in lines if line.strip()) with open(output_file, 'w') as out_file: for line in unique_lines: out_file.write(f"{line}\n") print("Unique lines written to output.txt successfully.") except FileNotFoundError: print("Error: input.txt not found.") except PermissionError: print("Error: Permission denied while opening the file.") except Exception as e: print(f"An error occurred: {e}") # Usage remove_duplicates('input.txt', 'output.txt')
问题排查
- 集合去重逻辑本身无错误:
set(line.strip()...)会自动对strip后的字符串去重,理论上不会出现重复行。出现重复的核心原因大概率是:input.txt中存在看似相同但实际字符不同的行(比如全角/半角符号、大小写差异、不可见空白字符、中文标点等),这类行strip后仍会被集合判定为不同元素。 - 额外问题:使用集合会打乱原文件中行的首次出现顺序,若需要保留原顺序,集合无法满足需求。
修正方案
方案1:排查并处理字符差异
先添加代码查看每行strip后的实际内容,确认是否存在隐形差异:
def check_lines(input_file): try: with open(input_file, 'r') as in_file: lines = in_file.readlines() for idx, line in enumerate(lines, 1): stripped = line.strip() if stripped: print(f"第{idx}行strip后内容:{repr(stripped)}") except Exception as e: print(f"检查出错:{e}") check_lines('input.txt')
运行后会输出每行的原始字符串表示(比如'Python is fun.'和'Python is fun.'会显示差异),根据结果针对性修正input.txt或添加字符串统一处理逻辑(比如转小写、替换全角符号等)。
方案2:保留原顺序的去重代码
如果需要保留原文件中首次出现行的顺序,改用dict.fromkeys()(Python 3.7+ 字典默认有序)实现去重,同时保留顺序:
def remove_duplicates(input_file, output_file): try: with open(input_file, 'r') as in_file: lines = in_file.readlines() # 用dict.fromkeys保留首次出现顺序,自动去重 unique_stripped = dict.fromkeys(line.strip() for line in lines if line.strip()) unique_lines = list(unique_stripped.keys()) with open(output_file, 'w') as out_file: for line in unique_lines: out_file.write(f"{line}\n") print("去重后的内容已成功写入output.txt。") except FileNotFoundError: print("错误:未找到input.txt文件。") except PermissionError: print("错误:打开文件时权限被拒绝。") except Exception as e: print(f"发生错误:{e}") # 调用函数 remove_duplicates('input.txt', 'output.txt')
补充说明
- 若只需去重不关心顺序,原代码本身逻辑正确,重复行问题必是输入文本存在隐形字符差异,需用
repr()排查。 - 使用
dict.fromkeys()既能去重,又能严格保留原文件中每行首次出现的顺序,更符合多数场景需求。
内容的提问来源于stack exchange,提问作者NMS
相关产品推荐
相关产品推荐

