按多字段汇总Pandas DataFrame并合并为单列
Pandas DataFrame分组汇总并多字段合并为单列
实现步骤
假设原始DataFrame结构如下(示例数据):
import pandas as pd df = pd.DataFrame({ 'ID': [1,1,1,2,2,2], 'LayerName': ['SC','SC','BLD','SC','BLD','BLD'], 'Name': ['B','R','S','B','K','S'], 'Count': [5,5,3,7,4,2] })
1. 分组计算与字段整理
按ID和LayerName分组,统计Count总和,同时收集去重后的Name集合:
grouped = df.groupby(['ID', 'LayerName']).agg( total_count=('Count', 'sum'), name_set=('Name', lambda x: tuple(sorted(set(x)))) ).reset_index()
2. 格式化分组字符串
将每个分组的信息整理为需求指定的格式:
grouped['formatted'] = grouped.apply( lambda row: f"{row['total_count']} - {row['LayerName']} : ({','.join(row['name_set'])})", axis=1 )
3. 按ID聚合生成最终Output列
把同一ID下的所有格式化字符串用; 拼接,得到每个ID对应的单行汇总:
result = grouped.groupby('ID')['formatted'].agg(' ; '.join).reset_index(name='Output')
最终结果
运行后result输出示例:
ID Output 0 1 10 - SC : (B,R) ; 3 - BLD : (S) 1 2 7 - SC : (B) ; 6 - BLD : (K,S)
内容的提问来源于stack exchange,提问作者lloyd
相关产品推荐
相关产品推荐

