Pandas DataFrame基于Column A去重时如何保留每组首尾行并删除其余重复行
Pandas按分组保留首尾行实现方案
前置依赖导入
首先导入需要的依赖库:
import pandas as pd import numpy as np
示例DataFrame构造
你提供的示例数据构造代码如下:
df = pd.DataFrame({ 'Column A': [12,12,12, 15, 16, 141, 141, 141, 141], 'Column B':['Apple' ,'Apple' ,'Apple' , 'Red', 'Blue', 'Yellow', 'Yellow', 'Yellow', 'Yellow'], 'Column C':[100, 50, np.nan , 23 , np.nan , 199 , np.nan , 1,np.nan] })
核心实现代码
有两种常用实现方式,按需选择即可:
方式1:布尔索引筛选(性能更优,适合大数据量)
通过分组正序、倒序计数筛选首尾行:
# 构造筛选掩码:正序计数为0(首行) 或 倒序计数为0(尾行) mask = df.groupby('Column A').cumcount().eq(0) | df.groupby('Column A').cumcount(ascending=False).eq(0) result = df[mask]
方式2:首尾拼接(逻辑更直观)
分别取每个分组的首行、尾行后合并去重,再按原始索引排序:
grouped = df.groupby('Column A', as_index=False, group_keys=False) result = pd.concat([grouped.head(1), grouped.tail(1)]).drop_duplicates().sort_index()
输出验证
打印result即可得到你需要的预期结果:
| Column A | Column B | Column C | |
|---|---|---|---|
| 0 | 12 | Apple | 100 |
| 2 | 12 | Apple | nan |
| 3 | 15 | Red | 23 |
| 4 | 16 | Blue | nan |
| 5 | 141 | Yellow | 199 |
| 8 | 141 | Yellow | nan |
内容的提问来源于stack exchange,提问作者new_bee
相关产品推荐
相关产品推荐

