You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas按客户分组将多行Cost值合并为单行多列

问题描述

需将DataFrame中同一客户的所有Cost值合并到同一行,第一个Cost值保留在原Cost列,其余依次放入Cost1、Cost2列,无对应值填充NaN。

原始DataFrame

CustomerCost
Jane Doe100
Jane Doe300
Jane Doe10
John Doe300

期望结果

CustomerCostCost1Cost2
Jane Doe10030010
John Doe300NaNNaN

解决方案

用Pandas可以通过分组、添加组内序号再重塑数据的方式实现需求,代码示例如下:

import pandas as pd

# 创建原始数据
df = pd.DataFrame({
    'Customer': ['Jane Doe', 'Jane Doe', 'Jane Doe', 'John Doe'],
    'Cost': [100, 300, 10, 300]
})

# 为每个客户的Cost记录添加组内序号(从0开始)
df['group_idx'] = df.groupby('Customer').cumcount()

# 透视转换,将组内序号转为列
pivot_df = df.pivot(index='Customer', columns='group_idx', values='Cost')

# 重命名列,匹配需求的Cost、Cost1、Cost2格式
pivot_df.columns = ['Cost'] + [f'Cost{i}' for i in range(1, pivot_df.shape[1])]

# 重置索引,让Customer回到列的位置
final_df = pivot_df.reset_index()

print(final_df)

步骤解释

  1. 添加组内序号:cumcount()为每个客户的每一条记录生成递增的序号,用来区分同一个客户下的不同Cost值。
  2. 透视重塑:pivot()把纵向的序号转为横向的列,每个序号对应一个Cost值。
  3. 列名调整:将第0列保留为Cost,后续列依次命名为Cost1、Cost2,符合需求格式。
  4. 重置索引:把Customer从索引转换为普通列,得到最终的表格结构。

执行代码后输出的结果完全匹配期望效果,缺失的位置会自动填充NaN。

内容的提问来源于stack exchange,提问作者Ceci V

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.28 13:43:22