如何按Child列分组并保留Child1列原有顺序转为列表?
问题:分组聚合后保留原DataFrame的行顺序
先看原始的DataFrame:
import pandas as pd df1 = pd.DataFrame({'Parent': ['Stay home', "Stay home","Stay home", 'Go swimming', "Go swimming","Go swimming"], 'Child': ['Severe weather', "Severe weather", "Severe weather", 'Not Severe weather', "Not Severe weather", "Not Severe weather"], 'Child1': ["Extreme rainy", "Extreme windy", "Severe snow", "Sunny", "some windy", "No snow"] })
需求是按Child列分组,把Child1列的值转为列表,但希望保留原DataFrame中分组出现的顺序(即先Severe weather对应的列表,再Not Severe weather对应的列表)。
尝试的代码:
def cast_to_list(df, col): return df.groupby('Child')[col].apply(list).tolist() list5=cast_to_list(df1, 'Child1') list5
得到的结果顺序不符合预期:
[['Sunny', 'some windy', 'No snow'], ['Extreme rainy', 'Extreme windy', 'Severe snow']]
解决方法
问题出在groupby默认会对分组键进行排序,所以'Not Severe weather'因为字母顺序排在了前面。只要在groupby时添加sort=False参数,就能保留分组键在原DataFrame中第一次出现的顺序:
修改后的代码:
def cast_to_list(df, col): # 添加sort=False,关闭分组键的自动排序 return df.groupby('Child', sort=False)[col].apply(list).tolist() list5=cast_to_list(df1, 'Child1') list5
执行后就能得到预期结果:
[['Extreme rainy', 'Extreme windy', 'Severe snow'], ['Sunny', 'some windy', 'No snow']]
内容的提问来源于stack exchange,提问作者xavi
相关产品推荐
相关产品推荐

