You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现两文件对比:提取file1中file2不存在的用户名

提取file1中未在file2出现的用户名

文件格式说明

  • file1.txt:每行格式为username:password,示例内容:
admin:admin
admin:meunsm
admin:12345
  • file2.txt:每行格式为source ip:source port > destination ip:destination port username:password,示例内容:
192.168.0.114:1137   >   192.168.0.193:21 csanders:echo

需求

用Python仅对比用户名,提取file1.txt中未在file2.txt出现的用户名并保存到新文件;因文件可能有数十万行,需用for循环(后续要将用户名存入数据库表)。

解决方案

原参考代码直接对比整行内容,不符合只对比用户名的需求,且一次性读入大文件会占用过多内存。下面是适配的实现:

# 先提取file2里的所有用户名存入集合,集合查询效率远高于列表
file2_usernames = set()
with open('file2.txt', 'r', encoding='utf-8') as f2:
    for line in f2:
        line = line.strip()
        if not line:
            continue  # 跳过空行
        # 分割行内容,最后一段是username:password格式
        user_pass_segment = line.split()[-1]
        username = user_pass_segment.split(':')[0]
        file2_usernames.add(username)

# 遍历file1,筛选未在file2出现的用户名并写入新文件
with open('file1.txt', 'r', encoding='utf-8') as f1, open('unused_usernames.txt', 'w', encoding='utf-8') as out_file:
    for line in f1:
        line = line.strip()
        if not line:
            continue
        username = line.split(':')[0]
        if username not in file2_usernames:
            # 可选择只写入用户名,或保留整行,这里按原格式写入整行
            out_file.write(line + '\n')
            # 后续存数据库的逻辑可直接加在此处,无需修改循环结构

代码说明

  1. 先逐行处理file2,提取每行的用户名存入集合:集合的in操作是O(1)复杂度,面对大文件时查询效率远超列表。
  2. 逐行处理file1,提取用户名后判断是否不在file2的用户名集合中,符合条件则写入结果文件。
  3. 全程逐行读取,不会一次性加载整个文件到内存,适配数十万行的大文件场景。
  4. 后续要存入数据库的话,直接在筛选条件的代码块内添加数据库插入语句即可,无需调整现有循环逻辑。

内容的提问来源于stack exchange,提问作者AHSAN YAZDANI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 01:20:23