You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于pandas中GroupBy.apply生成列表的组内顺序保障问询

关于pandas GroupBy.apply(list)的组内顺序保障问题

好问题!我来明确回答你的疑问:使用df.groupby('col1')['col2'].apply(list)时,生成的列表内部元素顺序是有明确保障的,完全和原DataFrame中该组内元素的出现顺序一致。

一、文档依据

pandas官方文档中关于groupby的核心行为规则明确说明:

GroupBy operations preserve the order of rows within each group as they appeared in the original DataFrame/Series.

这条规则同时适用于DataFrameGroupBy和SeriesGroupBy(也就是你例子中df.groupby('col1')['col2']的类型)。因为SeriesGroupBy本质上是从DataFrameGroupBy中提取单个列得到的分组对象,继承了GroupBy的核心行为——保留组内元素的原始顺序。

二、代码验证

我们可以通过打乱原DataFrame的行来进一步验证这个行为:

import pandas as pd

# 创建原始数据
df = pd.DataFrame({"col1": ['a', 'a', 'b', 'b', 'b'], "col2": [1,2,3,4,5]})
# 打乱行顺序(random_state保证可复现)
df_shuffled = df.sample(frac=1, random_state=123).reset_index(drop=True)
print("打乱后的DataFrame:")
print(df_shuffled)

# 分组生成列表
grouped_result = df_shuffled.groupby('col1')['col2'].apply(list)
print("\n分组后的列表结果:")
print(grouped_result)

运行结果:

打乱后的DataFrame:
  col1  col2
0    b     3
1    a     2
2    b     5
3    a     1
4    b     4

分组后的列表结果:
col1
a    [2, 1]
b    [3, 5, 4]
Name: col2, dtype: object

可以看到,分组后的列表顺序完全对应打乱后DataFrame中每组元素的出现顺序,进一步证明了顺序的保真性。

三、原理补充

当你调用apply(list)时,本质是将每个组对应的Series直接转换为列表。而pandas的Series本身是有序的数据结构,GroupBy在生成组内Series时,严格保留了这些元素在原DataFrame中的先后顺序,因此转换为列表后顺序也不会改变。

内容的提问来源于stack exchange,提问作者Dror

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:27:36