You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python对比两个txt文件时set求差集丢失重复差异行如何解决

问题原因

你原来的代码使用set计算差集,而集合的核心特性是自动去重,仅会保留唯一元素。因此即便two.txt内有2个未在one.txt中出现的c,集合中只会留存1个c,最终输出自然只有1行。

解决方法

你可以先把one.txt的行存入集合做存在性校验,再逐行遍历two.txt收集不存在的行,完整保留出现次数:

# 读取one.txt内容去重,用于快速判断行是否存在
with open('one.txt', 'r') as f:
    one_unique_lines = set(f.readlines())

result = []
with open('two.txt', 'r') as f:
    for line in f:
        # 可按需保留/删除空行过滤逻辑
        if line == '\n':
            continue
        if line not in one_unique_lines:
            result.append(line)

# 写入结果文件
with open('difff.txt', 'w') as f:
    f.writelines(result)

运行上述代码后,difff.txt就会输出2行c,完全匹配你的需求。

扩展方案(按计数差输出)

如果后续需要按两个文件的行计数差输出,比如one.txt有5个b、two.txt有6个b时输出多出来的1个b,可以使用Counter实现更精准的差集统计:

from collections import Counter

with open('one.txt', 'r') as f:
    count_one = Counter(f.readlines())
with open('two.txt', 'r') as f:
    count_two = Counter(f.readlines())

result = []
for line, cnt in count_two.items():
    if line == '\n':
        continue
    diff_cnt = cnt - count_one.get(line, 0)
    if diff_cnt > 0:
        result.extend([line] * diff_cnt)

with open('difff.txt', 'w') as f:
    f.writelines(result)

内容的提问来源于stack exchange,提问作者Mark Neyman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 01:45:02