Pandas按org_id分组聚合列:合并唯一值报错解决
问题:按org_id分组合并唯一值时出现FutureWarning
原始DataFrame
org_id org_name category org_status created_on modified_on location_id loc_status street_x city country 0 ORG-100023310 advanceCOR GmbH Industry,Pharmaceutical company ACTIVE 2016-10-18T15:38:34.322+02:00 2022-11-02T08:23:13.989+01:00 LOC-100052061 ACTIVE Fraunhoferstrasse 9a, Martinsried Planegg Germany 1 ORG-100023310 advanceCOR GmbH Industry,Pharmaceutical company ACTIVE 2016-10-18T15:38:34.322+02:00 2022-11-02T08:23:13.989+01:00 LOC-100032442 ACTIVE Lochhamer Strasse 29a, Martinsried Planegg Germany
需求
按org_id列分组,将每列的唯一值用|分隔后输出到新DataFrame,预期输出:
org_id org_name category org_status created_on modified_on location_id loc_status street_x city country 0 ORG-100023310 advanceCOR GmbH Industry,Pharmaceutical company ACTIVE 2016-10-18T15:38:34.322+02:00 2022-11-02T08:23:13.989+01:00 LOC-100052061 | LOC-100032442 ACTIVE Fraunhoferstrasse 9a, Martinsried | Lochhamer Strasse 29a, Martinsried Planegg Germany
尝试的代码及报错
尝试以下代码时出现FutureWarning:
join_unique = lambda x: '|'.join(x.unique()) df2 = df.groupby(['org_id'], as_index=False).agg(join_unique)
报错信息:
FutureWarning: ['loc_status', 'street_x', 'city', 'country'] did not aggregate successfully. If any error is raised this will raise in a future version of pandas. Drop these columns/ops to avoid this warning. df2 = df.groupby(['org_id'], as_index=False).agg(join_unique)
解决方案
出现警告的核心原因是聚合函数在处理部分列时,可能存在元素类型非字符串的情况,导致join操作无法顺利执行。修改聚合函数,先将所有元素转为字符串再去重拼接即可解决:
import pandas as pd # 构造示例DataFrame(如果已有可跳过) data = { 'org_id': ['ORG-100023310', 'ORG-100023310'], 'org_name': ['advanceCOR GmbH', 'advanceCOR GmbH'], 'category': ['Industry,Pharmaceutical company', 'Industry,Pharmaceutical company'], 'org_status': ['ACTIVE', 'ACTIVE'], 'created_on': ['2016-10-18T15:38:34.322+02:00', '2016-10-18T15:38:34.322+02:00'], 'modified_on': ['2022-11-02T08:23:13.989+01:00', '2022-11-02T08:23:13.989+01:00'], 'location_id': ['LOC-100052061', 'LOC-100032442'], 'loc_status': ['ACTIVE', 'ACTIVE'], 'street_x': ['Fraunhoferstrasse 9a, Martinsried', 'Lochhamer Strasse 29a, Martinsried'], 'city': ['Planegg', 'Planegg'], 'country': ['Germany', 'Germany'] } df = pd.DataFrame(data) # 修改后的聚合函数 join_unique = lambda x: ' | '.join(map(str, x.unique())) # 执行分组聚合 df2 = df.groupby(['org_id'], as_index=False).agg(join_unique) print(df2)
这段代码会将每列的唯一值转为字符串后,用|分隔拼接,最终得到符合预期的结果,同时消除FutureWarning。
内容的提问来源于stack exchange,提问作者rshar
相关产品推荐
相关产品推荐

