You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现原始CSV文件与系统生成CSV文件的全量行比对,提取前者未出现于后者的行?

Alright, let's tackle this problem head-on! The issue with your current code is that it's comparing lines positionally (line 1 of File1 vs line 1 of File2, line 2 vs line 2, etc.), but you need a full comparison where every line in File1 is checked against all lines in File2. Here's how to fix this:

Solution 1: Pure Python (Lightweight, No External Libraries)

This approach is great for large files since it doesn't load the entire dataset into memory at once, and uses a set for fast lookups.

# First, load all lines from File2 into a set for quick checking
with open('file2.csv', 'r') as f2:
    # Use rstrip('\n') to ignore line endings, ensuring consistent comparisons
    file2_lines = set(line.rstrip('\n') for line in f2)

# Now iterate through File1 and write missing lines to comparison.csv
with open('file1.csv', 'r') as f1, open('comparison.csv', 'w') as output_file:
    missing_count = 0
    for line in f1:
        cleaned_line = line.rstrip('\n')
        if cleaned_line not in file2_lines:
            output_file.write(line)
            missing_count += 1

print(f"Done! Found {missing_count} lines in File1 that aren't present in File2. Saved to comparison.csv")

How this works:

  • We first read every line from File2 and store them in a set—sets have near-instant lookup times, which makes checking if a line exists much faster than looping through File2 every time.
  • We then go through each line in File1, clean up the line ending (to avoid mismatches from \n vs \r\n), and if it's not in our File2 set, we write it directly to the output file.

Solution 2: Using Pandas (If You Prefer DataFrame Operations)

Since you imported pandas in your original code, here's a pandas-based approach that's clean and works well for smaller to medium-sized files:

import pandas as pd

# Read the CSV files (use header=None since there are no column headers)
# sep=' ' handles space-separated values; engine='python' deals with multiple spaces if present
df1 = pd.read_csv('file1.csv', header=None, sep=' ', engine='python')
df2 = pd.read_csv('file2.csv', header=None, sep=' ', engine='python')

# Perform a left merge to find rows only present in df1 (File1)
merged_df = df1.merge(df2, how='left', indicator=True)
missing_rows = merged_df[merged_df['_merge'] == 'left_only'].drop(columns='_merge')

# Save the result to comparison.csv (no headers/index to match your original format)
missing_rows.to_csv('comparison.csv', header=False, index=False, sep=' ')

print(f"Done! Saved {len(missing_rows)} missing lines to comparison.csv")

How this works:

  • We read both files into DataFrames, specifying no headers and space as the separator.
  • Using merge with how='left' keeps all rows from File1, and adds an _merge column indicating where each row came from.
  • We filter for rows marked left_only (only present in File1), remove the _merge column, and save the result.

Important Notes:

  • Make sure the lines in both files are exactly identical (same number of spaces, same date format, no trailing spaces). If there are formatting inconsistencies, you'll need to clean the data first (e.g., using strip() on each field, standardizing dates).
  • For very large files, the pure Python solution is more memory-efficient, as it processes lines one at a time instead of loading the entire file into memory.

内容的提问来源于stack exchange,提问作者francis fernandes

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 19:12:49