如何在Pandas DataFrame中按ID合并列表的列表类型的PATH列?
问题:按ID合并DataFrame中PATH列的列表内容
原始数据(dataframe1)
| ID | PATH |
|---|---|
| ABC | [[orange, apple, kiwi, peach], [strawberry, orange, kiwi, peach]] |
| ABC | [[apple, plum, peach], [apple, pear, peach]] |
| BCD | [[blueberry, plum, peach], [pear, apple, peach]] |
| BCD | [[plum, apple, peach], [banana, raspberry, peach]] |
期望输出(dataframe2)
按ID分组后,将每个ID对应的所有PATH子列表合并为一个二维列表:
| ID | PATH |
|---|---|
| ABC | [[orange, apple, kiwi, peach], [strawberry, orange, kiwi, peach], [apple, plum, peach], [apple, pear, peach]] |
| BCD | [[blueberry, plum, peach], [pear, apple, peach], [plum, apple, peach], [banana, raspberry, peach]] |
当前错误实现
执行以下代码后,得到了三层嵌套列表的错误结果(dataframe3):
df2 = df1.groupby('id')['PATH'].apply(list)
错误输出(dataframe3):
| ID | PATH |
|---|---|
| ABC | [[[orange, apple, kiwi, peach], [strawberry, orange, kiwi, peach], [apple, plum, peach], [apple, pear, peach]]] |
| BCD | [[[blueberry, plum, peach], [pear, apple, peach], [plum, apple, peach], [banana, raspberry, peach]]] |
解决方法
问题根源是apply(list)直接把每行的二维列表打包进新列表,导致三层嵌套。需要把每个ID下的子列表展平一层后合并,以下两种方法都可以实现:
方法1:用itertools.chain展平合并
from itertools import chain # 展平每个ID下的所有子列表,再转为列表,最后还原索引列 df2 = df1.groupby('ID')['PATH'].apply(lambda x: list(chain.from_iterable(x))).reset_index()
方法2:用列表推导式展平合并
# 通过嵌套循环遍历每个子列表,收集到新列表中 df2 = df1.groupby('ID')['PATH'].apply(lambda x: [sub_list for item in x for sub_list in item]).reset_index()
注意点
加上reset_index()是为了把分组后的ID从索引还原成普通列,让输出结构和原始DataFrame一致。
内容的提问来源于stack exchange,提问作者Mercury_Sun
相关产品推荐
相关产品推荐

