You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于itertools的Python脚本需求:对比两CSV文件首列并输出异同行

Python Script Using itertools to Compare Two CSV Files by First Column

Got it, let's tackle this problem. You want to compare the first columns of two CSV files (each row has two columns: a label and a value), identify shared rows (same label in both files) and unique rows (labels only present in one file). Here's a Python script using itertools to do exactly that:

import csv
import itertools

def read_csv_to_dict(file_path):
    """Read CSV file (two columns) into a dictionary {label: value}"""
    data_dict = {}
    with open(file_path, 'r') as f:
        reader = csv.reader(f, delimiter=' ')
        for row in reader:
            # Split row into label and value (handle spaces in the label)
            label = ' '.join(row[:-1])
            value = row[-1]
            data_dict[label] = value
    return data_dict

# Load data from both CSV files
file1_data = read_csv_to_dict('file1.csv')
file2_data = read_csv_to_dict('file2.csv')

# Combine all labels from both files and sort (required for itertools.groupby)
all_labels = sorted(itertools.chain(file1_data.keys(), file2_data.keys()))

# Classify labels into shared, file1-only, file2-only
shared_labels = []
file1_unique = []
file2_unique = []

for label, group in itertools.groupby(all_labels):
    group_size = len(list(group))
    if group_size == 2:
        shared_labels.append(label)
    elif label in file1_data:
        file1_unique.append(label)
    else:
        file2_unique.append(label)

# Print the results
print("=== Shared Rows (present in both files) ===")
for label in shared_labels:
    print(f"{label}: file1={file1_data[label]}, file2={file2_data[label]}")

print("\n=== Unique to file1.csv ===")
for label in file1_unique:
    print(f"{label}: {file1_data[label]}")

print("\n=== Unique to file2.csv ===")
for label in file2_unique:
    print(f"{label}: {file2_data[label]}")

How This Works:

  • Reading CSV Data: The read_csv_to_dict function converts each CSV into a dictionary, mapping the first column (label) to its corresponding value. It handles labels with spaces by joining all elements except the final value in each row.
  • Combining Labels: itertools.chain merges the label lists from both files into a single iterable, making it easy to process all labels at once.
  • Grouping with itertools: We sort the labels first (since groupby only groups consecutive identical elements), then use groupby to cluster identical labels. The size of each group tells us if the label exists in both files (size=2) or only one.
  • Result Output: We split the labels into three categories and print them with their corresponding values for clarity.

Example Output:

Running this script with your sample files will produce:

=== Shared Rows (present in both files) ===
C(2)—C(1): file1=1.5183, file2=1.5052
C(3)—C(1): file1=1.4991, file2=1.5037
C(3)—C(2): file1=1.49, file2=1.505
C(2)—C(1)—C(3): file1=59.19, file2=60.3
C(3)—C(1)—H(5): file1=118.4, file2=116.99

=== Unique to file1.csv ===
O(4)—C(3): 1.4104
H(10)—O(4): 0.964

=== Unique to file2.csv ===
S(4)—C(3): 1.7976
H(10)—S(4): 1.3445
C(2)—C(1)—H(6): 117.68

内容的提问来源于stack exchange,提问作者ahmed redha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:54:19