You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何更高效地基于Pandas DataFrame分类列生成新列?

更优的Pandas DataFrame列转换实现方法

需求说明

现有一个包含三列的Pandas DataFrame,其中col1为分类类型(仅含两个类别)。需要为col1的每个类别创建独立列,并将col2和col3的对应值存入这些新列,新列命名规则为原列名.类别。

原始DataFrame

col1col2col3
1e01
2f02
1g03
2h04

期望转换后的DataFrame

col2.1col2.2col3.1col3.2
ef0102
gh0304

现有实现代码

import pandas as pd

data = [
{'col1': 1, 'col2': 'e', 'col3': '01'}, 
{'col1': 2, 'col2': 'f', 'col3': '02'},
{'col1': 1, 'col2': 'g', 'col3': '03'},
{'col1': 2, 'col2': 'h', 'col3': '04'}
]

df = pd.DataFrame(data)

# 创建目标格式的DataFrame
new_df = pd.DataFrame({
    'col2.1': df[df['col1'] == 1]['col2'].values,
    'col2.2': df[df['col1'] == 2]['col2'].values,
    'col3.1': df[df['col1'] == 1]['col3'].values,
    'col3.2': df[df['col1'] == 2]['col3'].values
})
print(new_df)

更优实现方法

可以通过添加分组索引+unstack的方式实现,无需手动逐个筛选列,扩展性更强(后续col1新增类别也无需修改代码),且完全避免使用聚合函数:

import pandas as pd

data = [
{'col1': 1, 'col2': 'e', 'col3': '01'}, 
{'col1': 2, 'col2': 'f', 'col3': '02'},
{'col1': 1, 'col2': 'g', 'col3': '03'},
{'col1': 2, 'col2': 'h', 'col3': '04'}
]

df = pd.DataFrame(data)

# 为每个col1类别下的行添加组内序号,确保行对应关系
df['group_idx'] = df.groupby('col1').cumcount()

# 重塑数据:将col1的类别转为列层级
new_df = df.set_index(['group_idx', 'col1']).unstack('col1')
# 合并列名层级,生成"原列名.类别"的命名格式
new_df.columns = [f'{col}.{cat}' for col, cat in new_df.columns]
# 移除临时添加的组索引,得到最终格式
new_df = new_df.reset_index(drop=True)

print(new_df)

方法说明

  1. groupby('col1').cumcount():给每个col1类别下的行分配组内序号,保证转换后同一组的行能对应到同一行;
  2. set_index(['group_idx', 'col1']).unstack('col1'):把col1的类别转为列的层级,自动完成按类别拆分列的操作;
  3. 列名拼接:将原列名和类别用.连接,完全匹配需求的命名规则;
  4. 重置索引:去掉临时添加的group_idx列,得到和期望一致的结果。

这种方法代码更简洁,扩展性更好,不用针对每个类别手动编写筛选逻辑。

内容的提问来源于stack exchange,提问作者ciranzo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 09:43:11