如何在追加DataFrame时按自定义规则排除重复数据?
Solution to Append df1 to df2 with Custom Duplicate Exclusion
Got it, let's tackle this problem where we need to append df1 to df2 but exclude specific duplicate rows based on your custom rule. Here's a step-by-step approach using pandas:
Step 1: Break Down the Exclusion Logic
We only need to exclude rows from df1 if:
- The row's
namematches any row in df2 - That matching df2 row has a
gamevalue ofcar_version_2
Since df1 only contains rows withgame = car_v1, our focus is on filtering out names that already exist in df2'scar_version_2group.
Step 2: Implement the Solution
First, let's recreate your sample data to test with:
import pandas as pd # Sample df1 data df1 = pd.DataFrame({ 'name': ['Jane', 'Jamie', 'Kevin'], 'age': [7, 6, 9], 'game': ['car_v1', 'car_v1', 'car_v1'], 'col_d': ['foo', 'bar', 'bar'] }) # Sample df2 data df2 = pd.DataFrame({ 'name': ['Dave', 'Kevin', 'Jill', 'Chris', 'Kevin'], 'age': [1, 9, 6, 3, 9], 'game': ['train game', 'plane game', 'plane game', 'car_version_2', 'car_version_2'], 'col_d': ['foo', 'bar', 'bar', 'foo', 'bar'] })
Now apply the exclusion and append operation:
# Get all names in df2 that have a game value of 'car_version_2' exclude_names = df2[df2['game'] == 'car_version_2']['name'].unique() # Filter df1 to keep only rows where name is NOT in the exclusion list filtered_df1 = df1[~df1['name'].isin(exclude_names)] # Append the filtered df1 to df2, and reset the index to avoid duplicate labels final_df = pd.concat([df2, filtered_df1], ignore_index=True) print(final_df)
Step 3: Check the Result
Running the code above will produce exactly your expected output:
name age game col_d 0 Dave 1 train game foo 1 Kevin 9 plane game bar 2 Jill 6 plane game bar 3 Chris 3 car_version_2 foo 4 Kevin 9 car_version_2 bar 5 Jane 7 car_v1 foo 6 Jamie 6 car_v1 bar
The key here is the ~ operator, which flips the boolean check from "is in the exclusion list" to "is NOT in the exclusion list"—this ensures we only keep df1 rows that don't have a matching car_version_2 entry in df2.
内容的提问来源于stack exchange,提问作者David
相关产品推荐
相关产品推荐

