You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组处理:按指定规则保留组内目标行数据

解决方案

构造输入数据

先导入pandas并创建测试数据集:

import pandas as pd

df = pd.DataFrame({
    'ID': [1, 1, 2, 3, 4, 4],
    'Fruit': ['Apples', 'Oranges', 'Strawberries', 'Blueberries', 'Blueberries', 'Apples']
})

自定义聚合逻辑

编写一个函数实现规则:如果组内同时存在Apples和Oranges,返回Apples;否则返回组内第一条记录的水果:

def select_fruit(group):
    has_apples = (group['Fruit'] == 'Apples').any()
    has_oranges = (group['Fruit'] == 'Oranges').any()
    if has_apples and has_oranges:
        return 'Apples'
    return group['Fruit'].iloc[0]

分组聚合

用groupby按ID分组,通过apply调用自定义函数完成聚合,最后重置索引得到长格式结果:

result = df.groupby('ID').apply(select_fruit).reset_index(name='Fruit')

验证结果

运行代码后输出如下,完全符合需求:

ID         Fruit
0   1         Apples
1   2  Strawberries
2   3   Blueberries
3   4   Blueberries

说明

  • 使用apply是因为需要基于整个分组的水果集合做判断,它能获取每组的完整数据,比单纯的agg更灵活。
  • 用any()检查组内是否存在目标水果,逻辑清晰且效率较高。

内容的提问来源于stack exchange,提问作者Charlotte Deng

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 05:25:20