You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现两文件源IP对比并提取唯一对应数据至新文件

问题场景

有两个格式不同的文件,详情如下:

file1.txt

格式为 source ip -> destination ip,内容:

1.1.1.1 -> 0.0.0.0
2.2.2.2 -> 0.0.0.0
3.3.3.3 -> 1.1.1.1
4.4.4.4 -> 2.2.2.2

file2.txt

格式为 source ip,内容:

5.5.5.5
6.6.6.6
2.2.2.2
1.1.1.1

现有代码问题

当前使用的test.py代码无法正确对比两个文件,会把file1.txt的全部内容写入新文件:

with open("file1.txt") as fp1, open("file2.txt") as fp2, open("newfile.txt", "w") as fp3:
    i = 0
    k = 0
    while True:
        try:
            if i == 0:
                # at first get line from both file
                l1 = next(fp1)
                l2 = next(fp2)
            # if both the line is equal get another line
            if l1 == l2:
                try:
                    l1 = next(fp1)
                except StopIteration:
                    break
                l2 = next(fp2)
            # if line are not equal then put l1 in new file
            else:
                fp3.write(l1)
                try:
                    l1 = next(fp1)
                except StopIteration:
                    break
            i += 1
        except StopIteration:
            k += 1
            if k == 2:
                break
        except Exception as e:
            print(e)
            break

需求

仅对比两个文件的source ip,将file1.txt中source ip未在file2.txt出现的整行数据写入newfile.txt,预期输出:

3.3.3.3 -> 1.1.1.1
4.4.4.4 -> 2.2.2.2

解决方案

思路

  1. 先读取file2.txt的所有source ip,存入集合(集合查询效率远高于列表)
  2. 遍历file1.txt的每一行,提取每行的source ip
  3. 判断该source ip是否不在集合中,若是则写入newfile.txt

修正后的代码

# 读取file2的所有source ip到集合,自动去重且查询高效
with open("file2.txt") as fp2:
    source_ips_in_file2 = {line.strip() for line in fp2}

# 处理file1,筛选符合条件的行写入新文件
with open("file1.txt") as fp1, open("newfile.txt", "w") as fp3:
    for line in fp1:
        cleaned_line = line.strip()
        if not cleaned_line:
            continue  # 跳过空行
        # 分割提取source ip
        source_ip = cleaned_line.split(" -> ")[0].strip()
        if source_ip not in source_ips_in_file2:
            fp3.write(cleaned_line + "\n")

代码说明

  • 集合存储file2的source ip,成员查询操作时间复杂度为O(1),处理大文件时优势明显
  • 通过split(" -> ")精准分割file1的每行内容,提取出source ip
  • 增加空行处理逻辑,避免因文件末尾空行导致的错误
  • 代码逻辑简洁直观,易维护且执行效率高

内容的提问来源于stack exchange,提问作者AHSAN YAZDANI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 06:35:16