You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

对比两个DataFrame并生成字段匹配布尔表/标记的技术需求

Alright, let's solve this DataFrame comparison problem with two clear output schemes as requested. I'll use pandas (the go-to library for this sort of work) and walk through each step with code examples and explanations.

First, let's initialize the two sample DataFrames we're working with—this makes testing our code straightforward:

import pandas as pd

# Create original DataFrame Df
data_df = {
    'Country': ['南非', '美国', '美国', '英国'],
    'Category': ['健康', '鞋类', '服装', '汽车'],
    'Brand': ['Proctor', 'Nike', 'Forever21', 'Rover']
}
Df = pd.DataFrame(data_df)

# Create comparison DataFrame Df1
data_df1 = {
    'Country': ['南非', '美国', '美国', '英国'],
    'Category': ['健康', '鞋类', '服装', '£!"4'],
    'Brand': ['Proctor', 'Nike', 'Forever21', '11111']
}
Df1 = pd.DataFrame(data_df1)

方案1:逐字段布尔匹配结果表

This scheme generates a DataFrame Df3 that shows boolean (True/False) results for each field's match status, plus an optional flag for full row matches. It's great for quick, structured validation.

Code Implementation

# Create boolean match columns for each individual field
match_columns = pd.DataFrame({
    'Country_Match': Df['Country'] == Df1['Country'],
    'Category_Match': Df['Category'] == Df1['Category'],
    'Brand_Match': Df['Brand'] == Df1['Brand'],
    'Full_Row_Match': Df.eq(Df1).all(axis=1)  # Checks if entire row matches
})

# Combine original data with match results for context
Df3 = pd.concat([Df.add_prefix('Original_'), Df1.add_prefix('Comparison_'), match_columns], axis=1)

print(Df3)

Sample Output

Original_Country Original_Category Original_Brand Comparison_Country Comparison_Category Comparison_Brand  Country_Match  Category_Match  Brand_Match  Full_Row_Match
0               南非                健康        Proctor                 南非                健康        Proctor            True             True          True             True
1               美国                鞋类           Nike                 美国                鞋类           Nike            True             True          True             True
2               美国                服装     Forever21                 美国                服装     Forever21            True             True          True             True
3               英国                汽车          Rover                 英国              £!"4          11111            True            False         False            False

方案2:差异汇总表

This format prioritizes readability by highlighting exactly which fields don't match, alongside the original values from both DataFrames. It's ideal for quickly identifying what differs between rows.

Code Implementation

# Combine original and comparison data into one table
Df3 = pd.concat([Df.add_prefix('Original_'), Df1.add_prefix('Comparison_')], axis=1)

# Create a column to list mismatched fields for each row
def identify_mismatches(row):
    mismatched = []
    if row['Original_Country'] != row['Comparison_Country']:
        mismatched.append('Country')
    if row['Original_Category'] != row['Comparison_Category']:
        mismatched.append('Category')
    if row['Original_Brand'] != row['Comparison_Brand']:
        mismatched.append('Brand')
    return ', '.join(mismatched) if mismatched else 'All fields match'

Df3['Mismatched_Fields'] = Df3.apply(identify_mismatches, axis=1)

print(Df3)

Sample Output

Original_Country Original_Category Original_Brand Comparison_Country Comparison_Category Comparison_Brand Mismatched_Fields
0               南非                健康        Proctor                 南非                健康        Proctor    All fields match
1               美国                鞋类           Nike                 美国                鞋类           Nike    All fields match
2               美国                服装     Forever21                 美国                服装     Forever21    All fields match
3               英国                汽车          Rover                 英国              £!"4          11111    Category, Brand

内容的提问来源于stack exchange,提问作者yacub elmi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:08:53