对比两个DataFrame并生成字段匹配布尔表/标记的技术需求
Alright, let's solve this DataFrame comparison problem with two clear output schemes as requested. I'll use pandas (the go-to library for this sort of work) and walk through each step with code examples and explanations.
First, let's initialize the two sample DataFrames we're working with—this makes testing our code straightforward:
import pandas as pd # Create original DataFrame Df data_df = { 'Country': ['南非', '美国', '美国', '英国'], 'Category': ['健康', '鞋类', '服装', '汽车'], 'Brand': ['Proctor', 'Nike', 'Forever21', 'Rover'] } Df = pd.DataFrame(data_df) # Create comparison DataFrame Df1 data_df1 = { 'Country': ['南非', '美国', '美国', '英国'], 'Category': ['健康', '鞋类', '服装', '£!"4'], 'Brand': ['Proctor', 'Nike', 'Forever21', '11111'] } Df1 = pd.DataFrame(data_df1)
方案1:逐字段布尔匹配结果表
This scheme generates a DataFrame Df3 that shows boolean (True/False) results for each field's match status, plus an optional flag for full row matches. It's great for quick, structured validation.
Code Implementation
# Create boolean match columns for each individual field match_columns = pd.DataFrame({ 'Country_Match': Df['Country'] == Df1['Country'], 'Category_Match': Df['Category'] == Df1['Category'], 'Brand_Match': Df['Brand'] == Df1['Brand'], 'Full_Row_Match': Df.eq(Df1).all(axis=1) # Checks if entire row matches }) # Combine original data with match results for context Df3 = pd.concat([Df.add_prefix('Original_'), Df1.add_prefix('Comparison_'), match_columns], axis=1) print(Df3)
Sample Output
Original_Country Original_Category Original_Brand Comparison_Country Comparison_Category Comparison_Brand Country_Match Category_Match Brand_Match Full_Row_Match 0 南非 健康 Proctor 南非 健康 Proctor True True True True 1 美国 鞋类 Nike 美国 鞋类 Nike True True True True 2 美国 服装 Forever21 美国 服装 Forever21 True True True True 3 英国 汽车 Rover 英国 £!"4 11111 True False False False
方案2:差异汇总表
This format prioritizes readability by highlighting exactly which fields don't match, alongside the original values from both DataFrames. It's ideal for quickly identifying what differs between rows.
Code Implementation
# Combine original and comparison data into one table Df3 = pd.concat([Df.add_prefix('Original_'), Df1.add_prefix('Comparison_')], axis=1) # Create a column to list mismatched fields for each row def identify_mismatches(row): mismatched = [] if row['Original_Country'] != row['Comparison_Country']: mismatched.append('Country') if row['Original_Category'] != row['Comparison_Category']: mismatched.append('Category') if row['Original_Brand'] != row['Comparison_Brand']: mismatched.append('Brand') return ', '.join(mismatched) if mismatched else 'All fields match' Df3['Mismatched_Fields'] = Df3.apply(identify_mismatches, axis=1) print(Df3)
Sample Output
Original_Country Original_Category Original_Brand Comparison_Country Comparison_Category Comparison_Brand Mismatched_Fields 0 南非 健康 Proctor 南非 健康 Proctor All fields match 1 美国 鞋类 Nike 美国 鞋类 Nike All fields match 2 美国 服装 Forever21 美国 服装 Forever21 All fields match 3 英国 汽车 Rover 英国 £!"4 11111 Category, Brand
内容的提问来源于stack exchange,提问作者yacub elmi

