You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas groupby与agg的结果误解及相关问题咨询

问题描述

现有如下结构的DataFrame:

ID1ID2location1location2degree
S5S10NiceParis1
S9S6NiceParis6
S9S10NiceParis6
S11S12RomeParis6
S14S11MarseilleRome6
S15S11Les BrégRome6
S11S16RomeParis6
S13S11ParisRome6
S7S8BatzNice1

需求是按degree、location1、location2分组,将ID1和ID2聚合为列表,执行以下代码后遇到两个问题:

df_merge = df_merge.groupby(['degree','location1','location2']).agg({'ID1': lambda x: x.tolist(),'ID2': lambda x: x.tolist()},axis=1)

疑问:

  1. 如何让location1与location2的无序组合(如Nice&Paris和Paris&Nice)合并为同一分组,而非分为两行;
  2. 分组后的degree、location1、location2为何与ID1、ID2不在同一行,同时无法理解DataFrame转列表后的结构。

解决方案

1. 处理无序地点分组

核心思路是生成标准化的地点对键,将无序的地点组合转化为统一标识,以此作为分组依据:

import pandas as pd

# 生成排序后的地点元组作为分组键(确保Nice&Paris和Paris&Nice生成相同的键)
df_merge['location_pair'] = df_merge.apply(lambda row: tuple(sorted([row['location1'], row['location2']])), axis=1)

# 按degree和标准化地点键分组,聚合ID1、ID2为列表,同时重置索引将分组列转为普通列
df_result = df_merge.groupby(['degree', 'location_pair']).agg(
    ID1=('ID1', list),
    ID2=('ID2', list)
).reset_index()

# 可选:将标准化地点键拆分为两列,替换原location列
df_result[['location1', 'location2']] = pd.DataFrame(df_result['location_pair'].tolist(), index=df_result.index)
df_result = df_result.drop('location_pair', axis=1)

通过排序生成统一的地点对键,所有无序的地点组合会被归为同一分组。

2. 分组列与聚合列的显示问题及结构说明

  • 不在同一行的原因:groupby默认会将分组列设置为DataFrame的索引,所以degree、location相关列会显示在索引区域,而ID1、ID2是数据列。使用reset_index()方法可将索引还原为普通列,让所有列显示在同一行。
  • 聚合后列表结构说明:ID1和ID2列的每个单元格都是Python列表,包含对应分组下所有的ID1/ID2值。比如degree=6、地点对为(Paris, Rome)的分组,ID1会是['S11', 'S11', 'S13'],ID2会是['S12', 'S16', 'S11'],对应原表中所有Paris和Rome组合的行数据。

内容的提问来源于stack exchange,提问作者Pierre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 06:50:50