如何基于Series名称合并两组分组后的Pandas Series?
合并两组按name分组的Pandas Series并格式化内容
问题说明
现有两组通过以下代码生成的、按name字段分组的Pandas Object Series:
beforeseries = dfbefore.groupby('name', dropna=True)['order'].apply(list) print(beforeseries)
afterseries = dfafter.groupby('name', dropna=True)['order'].apply(list) print(afterseries)
它们的输出分别为:
beforeseries:
Name1 [first, second, third] Name2 [first, second, third] Name_n [first, second, third, fourth]
afterseries:
Name1 [fourth, fifth] Name2 [fourth, fifth, sixth] Name_n [fifth, sixth]
需要将二者合并,得到如下格式的结果:
Name1 ['first second third', 'fourth fifth'] Name2 ['first second third', 'fourth fifth sixth'] Name_n ['first second third fourth', 'fifth sixth']
解决方案
通过以下三步即可实现需求:
- 将每个Series中的列表元素用空格拼接为单个字符串
- 对齐两个Series的
name索引,合并为DataFrame - 将每行的两个字符串组合成列表,生成最终Series
具体代码
import pandas as pd # 1. 把列表转为空格分隔的字符串 before_str = beforeseries.apply(lambda x: ' '.join(x)) after_str = afterseries.apply(lambda x: ' '.join(x)) # 2. 按name对齐并合并为DataFrame merged_df = pd.DataFrame({'before': before_str, 'after': after_str}) # 3. 将每行的两个字段转为列表,生成目标Series result = merged_df.apply(lambda row: [row['before'], row['after']], axis=1) print(result)
输出结果
Name1 [first second third, fourth fifth] Name2 [first second third, fourth fifth sixth] Name_n [first second third fourth, fifth sixth]
内容的提问来源于stack exchange,提问作者JTA1618
相关产品推荐
相关产品推荐

