如何按Parent1列分组,将Child列值转为列表并仅组首行显示?
Pandas分组后仅在组首行显示列表
原始数据
import pandas as pd df1 = pd.DataFrame({ 'Parent': ['Stay home', "Stay home","Stay home", 'Go swimming', "Go swimming","Go swimming"], 'Parent1': ['Severe weather', "Severe weather", "Severe weather", 'Not Severe weather', "Not Severe weather", "Not Severe weather"], 'Child': ["Extreme rainy", "Extreme windy", "Severe snow", "Sunny", "some windy", "No snow"] })
需求说明
按Parent1列分组,将每组的Child列值合并为列表,仅在每组的第一行显示该列表,其余行留空,预期结果如下:
| Parent | Parent1 | Child | List1 | |
|---|---|---|---|---|
| 0 | Stay home | Severe weather | Extreme rainy | [Extreme rainy, Extreme windy, Severe snow] |
| 1 | Stay home | Severe weather | Extreme windy | |
| 2 | Stay home | Severe weather | Severe snow | |
| 3 | Go swimming | Not Severe weather | Sunny | [Sunny, some windy, No snow] |
| 4 | Go swimming | Not Severe weather | some windy | |
| 5 | Go swimming | Not Severe weather | No snow |
问题分析
你尝试的transform方法会给每组所有行填充相同列表,无法实现仅组首行显示的需求:
df1["list1"] = df1.groupby('Parent1')['Child'].transform(lambda x: x.tolist())
解决方案
方法一:分步实现
# 1. 计算每组的Child列表,得到以Parent1为索引的Series group_child_lists = df1.groupby('Parent1')['Child'].apply(list) # 2. 初始化List1列为空字符串 df1['List1'] = '' # 3. 获取每组第一行的索引,填充对应列表 first_row_indices = df1.groupby('Parent1').head(1).index df1.loc[first_row_indices, 'List1'] = df1.loc[first_row_indices, 'Parent1'].map(group_child_lists)
方法二:简洁写法
df1['List1'] = '' # 定位组首行并赋值 group_lists = df1.groupby('Parent1')['Child'].apply(list) df1.loc[df1.groupby('Parent1').head(1).index, 'List1'] = df1['Parent1'].map(group_lists).drop_duplicates()
代码说明
df1.groupby('Parent1').head(1).index:获取每组第一行的索引位置;groupby('Parent1')['Child'].apply(list):生成每组的Child值列表,索引为Parent1的唯一取值;- 通过
map将列表匹配到对应组的首行,其余行保持初始的空值状态。
内容的提问来源于stack exchange,提问作者xavi
相关产品推荐
相关产品推荐

