You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对比两个含员工ID与日期的文件,找出仅存在于单文件的记录

双文件记录差集提取方案

首先给出你提供的两个文件内容示例:

# file1 内容
empid1: date 1
empid1: date 2
empid2: date 3
empid2: date 4

# file2 内容
empid1: date 1
empid2: date 3
empid3: date 5

以下是两种常用实现方式:

方法1:Linux命令行快速实现(comm命令)

comm命令可以直接对比两个有序文件的差集,使用前需要先对文件排序去重:

  • 第一步生成排序去重后的临时文件:
    sort file1 | uniq > file1.sorted
    sort file2 | uniq > file2.sorted
    
  • 提取仅存在于file1的记录:
    comm -23 file1.sorted file2.sorted
    
    输出结果:
    empid1: date 2
    empid2: date 4
    
  • 提取仅存在于file2的记录:
    comm -13 file1.sorted file2.sorted
    
    输出结果:
    empid3: date 5
    

参数说明:comm的三个输出列默认分别是「仅第一个文件独有」、「仅第二个文件独有」、「两个文件共有」,-23就是隐藏第2、3列只保留第一列,-13就是隐藏第1、3列只保留第二列。


方法2:Python跨平台实现

如果需要在Windows等非Linux环境运行,可以用Python脚本处理,代码如下:

# 读取两个文件内容存入集合自动去重,跳过空行
with open('file1', 'r', encoding='utf-8') as f1:
    set1 = set(line.strip() for line in f1 if line.strip())
with open('file2', 'r', encoding='utf-8') as f2:
    set2 = set(line.strip() for line in f2 if line.strip())

# 计算两个方向的差集
only_file1 = set1 - set2
only_file2 = set2 - set1

# 输出结果
print("仅在file1中存在的记录:")
for line in only_file1:
    print(line)
print("\n仅在file2中存在的记录:")
for line in only_file2:
    print(line)

运行后输出和命令行方法完全一致。


内容的提问来源于stack exchange,提问作者NewBond007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 06:36:07