Pandas表格重塑需求:列出所有项并将空值填充为NaN或0
Pandas DataFrame 按指定格式重塑的实现方法
问题详情
初始代码
import pandas as pd df = pd.DataFrame({'race': ['one', 'one', 'one', 'two', 'two', 'two'], 'type': ['D', 'K', 'G', 'D', 'D', 'K'], 'item': ['x', 'y', 'z', 'q', 'x', 'y'], 'level': [1, 2, 1, 6, 2, 3]})
原始DataFrame输出
race type item level 0 one D x 1 1 one K y 2 2 one G z 1 3 two D q 6 4 two D x 2 5 two K y 3
期望的重塑格式
D K G item level item level item level race one x 1 y 2 z 1 two q 6 y 3 NaN NaN two x 2 NaN NaN NaN NaN
核心需求
- 仅调整格式便于阅读,无需数据聚合
- 单个
race内item唯一,不同race可重复 race需根据组内项数扩展行(如two有2个D类型项则显示2行)- 无对应项时填充
NaN或0
尝试过的代码及错误
使用pivot时触发重复索引错误:
df.pivot(index='race', columns='type', values=['level', 'item'])
错误信息:
ValueError: Index contains duplicate entries, cannot reshape
解决方案
问题根源是race作为单一索引存在重复项(如two对应多个D类型行),需添加组内序号作为多级索引的一部分,避免重复后再进行重塑。
完整实现代码
import pandas as pd df = pd.DataFrame({'race': ['one', 'one', 'one', 'two', 'two', 'two'], 'type': ['D', 'K', 'G', 'D', 'D', 'K'], 'item': ['x', 'y', 'z', 'q', 'x', 'y'], 'level': [1, 2, 1, 6, 2, 3]}) # 1. 为每个race+type组内的行添加序号,解决索引重复问题 df['seq'] = df.groupby(['race', 'type']).cumcount() # 2. 设置多级索引并按type展开列 result = df.set_index(['race', 'seq']).unstack('type')[['item', 'level']] # 3. 调整列顺序,匹配期望格式(先D/K/G,每个类型下先item后level) result = result.reorder_levels([1, 0], axis=1).sort_index(axis=1, level=0) # 4. 移除辅助序号索引,让race重复显示 result = result.reset_index(level='seq', drop=True) print(result)
输出结果
D K G item level item level item level race one x 1 y 2 z 1 two q 6 y 3 NaN NaN two x 2 NaN NaN NaN NaN
关键逻辑说明
groupby(['race', 'type']).cumcount():为每个race下的同类型数据生成从0开始的序号,确保race+seq组合索引唯一unstack('type'):将type列转换为列级别,实现横向展开- 列顺序调整:通过
reorder_levels和sort_index让列结构与期望格式完全匹配 - 移除
seq索引:让race按需求重复显示,符合阅读习惯
内容的提问来源于stack exchange,提问作者edge-case
相关产品推荐
相关产品推荐

