You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Pandas按条件筛选重复行并拆分DataFrame

问题与解决方案

需求描述

我有一个简单的DataFrame,需要按以下规则拆分:

  • 将Car列存在重复的行(排除Cond列为"X"的行)单独提取到一个DataFrame;
  • 主DataFrame中每个Car仅保留一行,优先保留Cond列为"X"的行。

我已经写了部分代码,但不知道如何让主DataFrame仅保留符合要求的行,有没有能同时完成重复行筛选拆分与主DataFrame处理的方法?

原始数据

CarYearSpeedCond
BMW2001150X
BMW2000150
Audi1997200
Audi2000200
Audi2012200X
Fiat2020180
Mazda2022183

现有代码

import pandas as pd
import numpy as np

cars = {'Car': {0: 'BMW', 1: 'BMW', 2: 'Audi', 3: 'Audi', 4: 'Audi', 5: 'Fiat', 6: 'Mazda'},
        'Year': {0: 2001, 1: 2000, 2: 1997, 3: 2000, 4: 2012, 5: 2020, 6: 2022},
        'Speed': {0: 150, 1: 150, 2: 200, 3: 200, 4: 200, 5: 180, 6: 183},
        'Cond': {0: 'X', 1: np.nan, 2: 'X', 3: np.nan, 4: np.nan, 5: np.nan, 6: np.nan}}

df = pd.DataFrame.from_dict(cars)
df_duplicates = df.loc[df.duplicated(subset=['Car'], keep = False)].loc[df['Cond']!='X']

解决方案

可以通过标记优先级+分组筛选的方式同时完成两个需求,具体代码如下:

import pandas as pd
import numpy as np

cars = {'Car': {0: 'BMW', 1: 'BMW', 2: 'Audi', 3: 'Audi', 4: 'Audi', 5: 'Fiat', 6: 'Mazda'},
        'Year': {0: 2001, 1: 2000, 2: 1997, 3: 2000, 4: 2012, 5: 2020, 6: 2022},
        'Speed': {0: 150, 1: 150, 2: 200, 3: 200, 4: 200, 5: 180, 6: 183},
        'Cond': {0: 'X', 1: np.nan, 2: 'X', 3: np.nan, 4: np.nan, 5: np.nan, 6: np.nan}}

df = pd.DataFrame.from_dict(cars)

# 给行设置优先级:Cond为"X"的行优先级为0(最高),其他为1
df['priority'] = np.where(df['Cond'] == 'X', 0, 1)

# 按优先级排序后,每个Car组取第一行,得到主DataFrame
df_main = df.sort_values('priority').groupby('Car').first().reset_index()
# 删除临时的priority列
df_main = df_main.drop('priority', axis=1)

# 提取需要拆分的重复行:原始数据中不在主DataFrame里的行
df_duplicates = df[~df.index.isin(df_main.index)]

# 输出结果验证
print("主DataFrame(每个Car仅保留一行,优先Cond为X):")
print(df_main)
print("\n拆分出的重复行(Car重复且Cond不为X):")
print(df_duplicates)

代码说明

  1. 优先级标记:用np.where给Cond为"X"的行标记最高优先级,确保分组时优先被选中;
  2. 分组筛选主DataFrame:先按优先级排序,再按Car分组取第一行,保证每个Car只保留符合要求的行;
  3. 拆分重复行:通过索引对比,筛选出原始数据中未被保留在主DataFrame里的行,就是需要单独提取的重复行。

这种方法逻辑清晰,能一次性完成两个需求,避免多次筛选的冗余操作。


内容的提问来源于stack exchange,提问作者AnnAc0nda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 13:30:52