Pandas基于多列提取前n大值对应行的简便实现方法咨询
实现方案
直接对数值列批量处理即可,不需要单独拆分每列再合并,代码如下:
import pandas as pd # 原始DataFrame df = pd.DataFrame({ 'A': [1, 0.7, 0, 0.5, 0.3, 0.3], 'B': [0.6, 0.1, 0.4, 0.3, 0.9, 0.3], 'C': [0.6, 0.3, 0.6, 0.8, 0.9, 0.5], 'ID': ['a', 'b', 'c', 'd', 'e', 'f'] }) n = 2 res = ( df.set_index('ID') # 对A/B/C每一列,非前n大的数值替换为NaN .apply(lambda col: col.where(col.isin(col.nlargest(n)))) # 过滤掉三列全为NaN的无效行 .dropna(how='all') # 可选:将ID从索引恢复为普通列 .reset_index() )
输出结果(n=2时)
| ID | A | B | C |
|---|---|---|---|
| a | 1.0 | 0.6 | NaN |
| b | 0.7 | NaN | NaN |
| d | NaN | NaN | 0.8 |
| e | NaN | 0.9 | 0.9 |
并列值处理说明
如果某列存在多个和第n大数值相等的并列项,需要全部保留的话,可以把列处理逻辑替换为rank判断,适配并列场景:
.apply(lambda col: col.where(col.rank(ascending=False, method='min') <= n))
内容的提问来源于stack exchange,提问作者Sam
相关产品推荐
相关产品推荐

