Pandas groupby聚合非数值数据:按sid分组取cat_type首个非na值
实现方法
首先处理数据中的空值标识:如果你的cat_type列里的na是字符串,需要先替换为pandas可识别的空值类型:
import pandas as pd # 替换字符串na为 pandas 原生空值 df['cat_type'] = df['cat_type'].replace('na', pd.NA)
之后直接在agg方法中使用pandas内置的first聚合函数即可,该函数会自动跳过分组内的空值,返回第一个非空的取值:
df_groupby = df.groupby(['sid'], as_index=False).agg( score_max = ('score','max'), cat_type_first_row = ('cat_type', 'first') )
如果需要兼容旧版本pandas,或者需要自定义全部分组都为空时的返回值,可以用自定义聚合逻辑:
df_groupby = df.groupby(['sid'], as_index=False).agg( score_max = ('score','max'), # 逻辑为:先去掉分组内的空值,取第一个值,分组全空则返回空值 cat_type_first_row = ('cat_type', lambda x: x.dropna().iloc[0] if not x.dropna().empty else pd.NA) )
执行上述代码后得到的结果和你预期的输出完全一致。
内容的提问来源于stack exchange,提问作者floss
相关产品推荐
相关产品推荐

