You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用set与strip()处理文本文件去重失效,请求排查修正代码

文本去重程序重复行问题排查与修正

问题场景

需求:使用Python上下文管理器读取input.txt,忽略首尾空白判定重复行,去重后写入output.txt并处理异常。

示例input.txt内容:

Hello, World!
Python is fun.
Python is fun.
Python is fun.

期望output.txt结果:

Hello, World!
Python is fun.

但运行代码后,实际output.txt出现重复行:

Hello, World!
Python is fun.
Python is fun.

用户提供的代码:

def remove_duplicates(input_file, output_file):
    try:
        with open(input_file, 'r') as in_file:
            lines = in_file.readlines()

        unique_lines = set(line.strip() for line in lines if line.strip())

        with open(output_file, 'w') as out_file:
            for line in unique_lines:
                out_file.write(f"{line}\n")

        print("Unique lines written to output.txt successfully.")
    except FileNotFoundError:
        print("Error: input.txt not found.")
    except PermissionError:
        print("Error: Permission denied while opening the file.")
    except Exception as e:
        print(f"An error occurred: {e}")

# Usage
remove_duplicates('input.txt', 'output.txt')

问题排查

  1. 集合去重逻辑本身无错误:set(line.strip()...)会自动对strip后的字符串去重,理论上不会出现重复行。出现重复的核心原因大概率是:input.txt中存在看似相同但实际字符不同的行(比如全角/半角符号、大小写差异、不可见空白字符、中文标点等),这类行strip后仍会被集合判定为不同元素。
  2. 额外问题:使用集合会打乱原文件中行的首次出现顺序,若需要保留原顺序,集合无法满足需求。

修正方案

方案1:排查并处理字符差异

先添加代码查看每行strip后的实际内容,确认是否存在隐形差异:

def check_lines(input_file):
    try:
        with open(input_file, 'r') as in_file:
            lines = in_file.readlines()
        for idx, line in enumerate(lines, 1):
            stripped = line.strip()
            if stripped:
                print(f"第{idx}行strip后内容:{repr(stripped)}")
    except Exception as e:
        print(f"检查出错:{e}")

check_lines('input.txt')

运行后会输出每行的原始字符串表示(比如'Python is fun.'和'Python is fun.'会显示差异),根据结果针对性修正input.txt或添加字符串统一处理逻辑(比如转小写、替换全角符号等)。

方案2:保留原顺序的去重代码

如果需要保留原文件中首次出现行的顺序,改用dict.fromkeys()(Python 3.7+ 字典默认有序)实现去重,同时保留顺序:

def remove_duplicates(input_file, output_file):
    try:
        with open(input_file, 'r') as in_file:
            lines = in_file.readlines()

        # 用dict.fromkeys保留首次出现顺序,自动去重
        unique_stripped = dict.fromkeys(line.strip() for line in lines if line.strip())
        unique_lines = list(unique_stripped.keys())

        with open(output_file, 'w') as out_file:
            for line in unique_lines:
                out_file.write(f"{line}\n")

        print("去重后的内容已成功写入output.txt。")
    except FileNotFoundError:
        print("错误:未找到input.txt文件。")
    except PermissionError:
        print("错误:打开文件时权限被拒绝。")
    except Exception as e:
        print(f"发生错误:{e}")

# 调用函数
remove_duplicates('input.txt', 'output.txt')

补充说明

  • 若只需去重不关心顺序,原代码本身逻辑正确,重复行问题必是输入文本存在隐形字符差异,需用repr()排查。
  • 使用dict.fromkeys()既能去重,又能严格保留原文件中每行首次出现的顺序,更符合多数场景需求。

内容的提问来源于stack exchange,提问作者NMS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 05:53:23