You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas表格重塑需求:列出所有项并将空值填充为NaN或0

Pandas DataFrame 按指定格式重塑的实现方法

问题详情

初始代码

import pandas as pd

df = pd.DataFrame({'race': ['one', 'one', 'one', 'two', 'two', 'two'],
                'type': ['D', 'K', 'G', 'D', 'D', 'K'],
                'item': ['x', 'y', 'z', 'q', 'x', 'y'],
                'level': [1, 2, 1, 6, 2, 3]})

原始DataFrame输出

race    type    item    level
0   one     D       x       1
1   one     K       y       2
2   one     G       z       1
3   two     D       q       6
4   two     D       x       2
5   two     K       y       3

期望的重塑格式

D               K               G   
        item    level   item    level   item    level
race
one     x       1       y       2       z       1
two     q       6       y       3       NaN     NaN
two     x       2       NaN     NaN     NaN     NaN

核心需求

  • 仅调整格式便于阅读,无需数据聚合
  • 单个race内item唯一,不同race可重复
  • race需根据组内项数扩展行(如two有2个D类型项则显示2行)
  • 无对应项时填充NaN或0

尝试过的代码及错误

使用pivot时触发重复索引错误:

df.pivot(index='race', columns='type', values=['level', 'item'])

错误信息:

ValueError: Index contains duplicate entries, cannot reshape

解决方案

问题根源是race作为单一索引存在重复项(如two对应多个D类型行),需添加组内序号作为多级索引的一部分,避免重复后再进行重塑。

完整实现代码

import pandas as pd

df = pd.DataFrame({'race': ['one', 'one', 'one', 'two', 'two', 'two'],
                'type': ['D', 'K', 'G', 'D', 'D', 'K'],
                'item': ['x', 'y', 'z', 'q', 'x', 'y'],
                'level': [1, 2, 1, 6, 2, 3]})

# 1. 为每个race+type组内的行添加序号,解决索引重复问题
df['seq'] = df.groupby(['race', 'type']).cumcount()

# 2. 设置多级索引并按type展开列
result = df.set_index(['race', 'seq']).unstack('type')[['item', 'level']]

# 3. 调整列顺序,匹配期望格式(先D/K/G,每个类型下先item后level)
result = result.reorder_levels([1, 0], axis=1).sort_index(axis=1, level=0)

# 4. 移除辅助序号索引,让race重复显示
result = result.reset_index(level='seq', drop=True)

print(result)

输出结果

D               K               G   
        item    level   item    level   item    level
race
one     x       1       y       2       z       1
two     q       6       y       3       NaN     NaN
two     x       2       NaN     NaN     NaN     NaN

关键逻辑说明

  • groupby(['race', 'type']).cumcount():为每个race下的同类型数据生成从0开始的序号,确保race+seq组合索引唯一
  • unstack('type'):将type列转换为列级别,实现横向展开
  • 列顺序调整:通过reorder_levels和sort_index让列结构与期望格式完全匹配
  • 移除seq索引:让race按需求重复显示,符合阅读习惯

内容的提问来源于stack exchange,提问作者edge-case

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 05:34:54