基于itertools的Python脚本需求:对比两CSV文件首列并输出异同行
Python Script Using itertools to Compare Two CSV Files by First Column
Got it, let's tackle this problem. You want to compare the first columns of two CSV files (each row has two columns: a label and a value), identify shared rows (same label in both files) and unique rows (labels only present in one file). Here's a Python script using itertools to do exactly that:
import csv import itertools def read_csv_to_dict(file_path): """Read CSV file (two columns) into a dictionary {label: value}""" data_dict = {} with open(file_path, 'r') as f: reader = csv.reader(f, delimiter=' ') for row in reader: # Split row into label and value (handle spaces in the label) label = ' '.join(row[:-1]) value = row[-1] data_dict[label] = value return data_dict # Load data from both CSV files file1_data = read_csv_to_dict('file1.csv') file2_data = read_csv_to_dict('file2.csv') # Combine all labels from both files and sort (required for itertools.groupby) all_labels = sorted(itertools.chain(file1_data.keys(), file2_data.keys())) # Classify labels into shared, file1-only, file2-only shared_labels = [] file1_unique = [] file2_unique = [] for label, group in itertools.groupby(all_labels): group_size = len(list(group)) if group_size == 2: shared_labels.append(label) elif label in file1_data: file1_unique.append(label) else: file2_unique.append(label) # Print the results print("=== Shared Rows (present in both files) ===") for label in shared_labels: print(f"{label}: file1={file1_data[label]}, file2={file2_data[label]}") print("\n=== Unique to file1.csv ===") for label in file1_unique: print(f"{label}: {file1_data[label]}") print("\n=== Unique to file2.csv ===") for label in file2_unique: print(f"{label}: {file2_data[label]}")
How This Works:
- Reading CSV Data: The
read_csv_to_dictfunction converts each CSV into a dictionary, mapping the first column (label) to its corresponding value. It handles labels with spaces by joining all elements except the final value in each row. - Combining Labels:
itertools.chainmerges the label lists from both files into a single iterable, making it easy to process all labels at once. - Grouping with itertools: We sort the labels first (since
groupbyonly groups consecutive identical elements), then usegroupbyto cluster identical labels. The size of each group tells us if the label exists in both files (size=2) or only one. - Result Output: We split the labels into three categories and print them with their corresponding values for clarity.
Example Output:
Running this script with your sample files will produce:
=== Shared Rows (present in both files) === C(2)—C(1): file1=1.5183, file2=1.5052 C(3)—C(1): file1=1.4991, file2=1.5037 C(3)—C(2): file1=1.49, file2=1.505 C(2)—C(1)—C(3): file1=59.19, file2=60.3 C(3)—C(1)—H(5): file1=118.4, file2=116.99 === Unique to file1.csv === O(4)—C(3): 1.4104 H(10)—O(4): 0.964 === Unique to file2.csv === S(4)—C(3): 1.7976 H(10)—S(4): 1.3445 C(2)—C(1)—H(6): 117.68
内容的提问来源于stack exchange,提问作者ahmed redha
相关产品推荐
相关产品推荐

