You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas按列分组统计元素出现次数并保留分组首行如何实现

问题描述

现有以下结构的DataFrame:

import pandas as pd
import numpy as np

df = pd.DataFrame({
    'Column A': [12,12,12, 13, 15, 16, 141, 141, 141, 141],
    'Column B':['Apple' ,'Apple' ,'Orange' ,'Apple' , np.nan, 'Orange', 'Apple', np.nan, 'Apple', 'Apple']
})
需求说明
  • 按Column A的值分组:
    1. 若分组内行数>1(存在重复分组):统计Column B中Orange的出现次数填入新列Column C,统计Apple的出现次数填入新列Column D
    2. 若分组内行数=1(非重复分组):若Column B为NaN,则Column C、Column D填NaN;否则按实际值匹配,是Orange则C=1、D=0,是Apple则C=0、D=1
  • 分组Column E规则:若分组内Column B存在Orange填Yes,不存在填No,若分组Column B全为NaN则填NaN
  • 最终输出每个分组仅保留第一行
实现代码
def process_group(group):
    # 取分组第一行作为基础返回行
    res = group.iloc[0:1].copy()
    b_values = group['Column B']
    
    # 处理Column C、Column D
    if b_values.isna().all():
        res['Column C'] = np.nan
        res['Column D'] = np.nan
    else:
        res['Column C'] = (b_values == 'Orange').sum()
        res['Column D'] = (b_values == 'Apple').sum()
    
    # 处理Column E
    if b_values.isna().all():
        res['Column E'] = np.nan
    else:
        res['Column E'] = 'Yes' if (b_values == 'Orange').any() else 'No'
    
    return res

# 按Column A分组处理后重置索引
result = df.groupby('Column A', as_index=False).apply(process_group).reset_index(drop=True)
# 可选:将C、D列转为支持NaN的整数类型,输出格式更贴合预期
result[['Column C', 'Column D']] = result[['Column C', 'Column D']].astype('Int64')

print(result)
输出结果

运行代码后得到的结果和预期完全一致:

Column A Column B  Column C  Column D Column E
0        12    Apple         1         2      Yes
1        13    Apple         0         1       No
2        15      NaN      <NA>      <NA>      NaN
3        16   Orange         1         0      Yes
4       141    Apple         0         3       No

内容的提问来源于stack exchange,提问作者new_bee

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 18:24:05