You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中基于行数据生成新列(Pandas数据框转换)

Pandas DataFrame 宽表转换实现代码

需求说明

现有如下长格式DataFrame:

ClientIdProductQuantity
01Apples2
01Oranges3
01Bananas1
02Apples4
02Bananas2

需要转换为宽格式,其中Product_xxx列为二元变量(客户购买过对应产品则为1,未购买则为0),最终结构如下:

ClientIdProduct_ApplesQuantity_ApplesProduct_OrangesQuantity_OrangesProduct_BananasQuantity_Bananas
01121311
02140012

实现代码

import pandas as pd

# 构造原始DataFrame(已有数据集可跳过此步)
df = pd.DataFrame({
    'ClientId': ['01', '01', '01', '02', '02'],
    'Product': ['Apples', 'Oranges', 'Bananas', 'Apples', 'Bananas'],
    'Quantity': [2, 3, 1, 4, 2]
})

# 添加标识列,用于生成二元变量
df['Product_Flag'] = 1

# 生成透视表,聚合数量和标识值,缺失值补0
pivot_result = df.pivot_table(
    index='ClientId',
    columns='Product',
    values=['Quantity', 'Product_Flag'],
    fill_value=0,
    aggfunc='max'
)

# 重命名列,调整为目标格式
pivot_result.columns = [f'{col_type}_{product}' for col_type, product in pivot_result.columns]

# 重置索引并调整列顺序
final_df = pivot_result.reset_index()[
    ['ClientId', 'Product_Apples', 'Quantity_Apples',
     'Product_Oranges', 'Quantity_Oranges',
     'Product_Bananas', 'Quantity_Bananas']
]

# 查看结果
print(final_df)

代码说明

  1. 构造原始数据:如果已有现成的DataFrame,直接替换即可,无需重复构造;
  2. 添加标识列:新增Product_Flag列固定为1,用于后续生成Product_xxx的二元标识;
  3. 透视表转换:按ClientId分组,Product作为列维度,分别聚合购买数量(取对应值)和标识(取最大值,确保存在则为1、缺失补0);
  4. 列名调整:将透视表的多级列名重命名为前缀_产品名的格式;
  5. 列顺序整理:按照需求的列顺序重新排列,得到最终结果。

内容的提问来源于stack exchange,提问作者wlog

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 23:02:38