You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何仅用groupby实现DataFrame分组后的多列存在性标记?

问题描述

给定如下结构的pandas DataFrame:

import pandas as pd
df = pd.DataFrame({'id':[1,2,3,1, 1], 'time_stamp_date':['12','12', '12', '14', '14'], 'sth':['col1','col1', 'col2','col2', 'col3']})

需求:针对sht_list中的每个列名,标记其在对应id和time_stamp_date分组下是否存在于sth列,生成col1、col2等标记列(存在为1,不存在为0)。

目前已有可行代码(借助transform实现):

df_out = df.assign(**{col: df.groupby(['id','time_stamp_date']).sth.transform(lambda x: (x==col).any()).astype(int) for col in sht_list})

但希望仅使用groupby(不借助transform及drop_duplicates),尝试的代码报错:

df_out = df.groupby(['id','time_stamp_date'])[['sth']].agg({lambda x: (x==col).any().astype(int) for col in sht_list})

错误信息:

SpecificationError: Function names must be unique, found multiple named <lambda>

请问能否修改该代码使其生效?


解决方案

报错原因是传给agg的字典将所有lambda函数作为键,导致键名重复(均为<lambda>),pandas无法区分各聚合操作对应的列名。只需将字典的键替换为sht_list中的列名,保留对应判断逻辑即可。

修改后的完整代码:

import pandas as pd

df = pd.DataFrame({'id':[1,2,3,1, 1], 'time_stamp_date':['12','12', '12', '14', '14'], 'sth':['col1','col1', 'col2','col2', 'col3']})
sht_list = ['col1', 'col2', 'col3']

df_out = df.groupby(['id','time_stamp_date'])['sth'].agg(
    {col: lambda x: int((x == col).any()) for col in sht_list}
).reset_index()

关键说明:

  1. 用sht_list的元素作为字典键,pandas会直接将这些值作为生成的标记列名,避免键重复问题。
  2. 把astype(int)替换为int(),实现相同的类型转换,写法更简洁。
  3. 调用reset_index()将分组的id和time_stamp_date从索引转回普通列,匹配需求的输出结构。

执行后得到的df_out结构:

idtime_stamp_datecol1col2col3
112100
114011
212100
312010

内容的提问来源于stack exchange,提问作者corianne1234

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 09:13:27