You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在追加DataFrame时按自定义规则排除重复数据?

Solution to Append df1 to df2 with Custom Duplicate Exclusion

Got it, let's tackle this problem where we need to append df1 to df2 but exclude specific duplicate rows based on your custom rule. Here's a step-by-step approach using pandas:

Step 1: Break Down the Exclusion Logic

We only need to exclude rows from df1 if:

  • The row's name matches any row in df2
  • That matching df2 row has a game value of car_version_2
    Since df1 only contains rows with game = car_v1, our focus is on filtering out names that already exist in df2's car_version_2 group.

Step 2: Implement the Solution

First, let's recreate your sample data to test with:

import pandas as pd

# Sample df1 data
df1 = pd.DataFrame({
    'name': ['Jane', 'Jamie', 'Kevin'],
    'age': [7, 6, 9],
    'game': ['car_v1', 'car_v1', 'car_v1'],
    'col_d': ['foo', 'bar', 'bar']
})

# Sample df2 data
df2 = pd.DataFrame({
    'name': ['Dave', 'Kevin', 'Jill', 'Chris', 'Kevin'],
    'age': [1, 9, 6, 3, 9],
    'game': ['train game', 'plane game', 'plane game', 'car_version_2', 'car_version_2'],
    'col_d': ['foo', 'bar', 'bar', 'foo', 'bar']
})

Now apply the exclusion and append operation:

# Get all names in df2 that have a game value of 'car_version_2'
exclude_names = df2[df2['game'] == 'car_version_2']['name'].unique()

# Filter df1 to keep only rows where name is NOT in the exclusion list
filtered_df1 = df1[~df1['name'].isin(exclude_names)]

# Append the filtered df1 to df2, and reset the index to avoid duplicate labels
final_df = pd.concat([df2, filtered_df1], ignore_index=True)

print(final_df)

Step 3: Check the Result

Running the code above will produce exactly your expected output:

name  age            game col_d
0   Dave    1      train game   foo
1  Kevin    9      plane game   bar
2   Jill    6      plane game   bar
3  Chris    3  car_version_2   foo
4  Kevin    9  car_version_2   bar
5   Jane    7          car_v1   foo
6  Jamie    6          car_v1   bar

The key here is the ~ operator, which flips the boolean check from "is in the exclusion list" to "is NOT in the exclusion list"—this ensures we only keep df1 rows that don't have a matching car_version_2 entry in df2.

内容的提问来源于stack exchange,提问作者David

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:50:07