You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Pandas按filename分组,基于area统计量筛选对应fidx?

按分组筛选具有分布代表性的行

实现步骤

针对你的需求,我们可以通过分组计算统计量+匹配最接近的行来完成,具体代码和说明如下:

核心代码

假设你的DataFrame变量名为df:

import pandas as pd

def pick_representative_rows(group):
    # 计算每组需要的关键统计值(分位数用就近匹配)
    target_values = [
        group['area'].min(),
        group['area'].max(),
        group['area'].quantile(0.33, interpolation='nearest'),
        group['area'].quantile(0.66, interpolation='nearest')
    ]
    
    # 筛选每组中最接近各统计值的行
    selected = []
    for val in target_values:
        # 找绝对差最小的行索引
        closest_row_idx = (group['area'] - val).abs().idxmin()
        selected.append(group.loc[closest_row_idx])
    
    # 去重(避免极端值重复选中)并返回
    return pd.DataFrame(selected).drop_duplicates()

# 分组执行并整理结果
final_df = df.groupby('filename', group_keys=False).apply(pick_representative_rows)

关键细节说明

  • quantile的interpolation='nearest'参数:直接取数据中最接近目标分位数的实际值,不用严格计算精确分位,符合你的需求
  • 绝对差匹配:确保选中的是数据集中真实存在的行,而不是统计计算出的虚拟值
  • drop_duplicates():防止极端情况下某一行同时是最小值和最大值(比如组内所有值相同),避免结果出现重复行

内容的提问来源于stack exchange,提问作者Kenan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 21:52:39