Pandas中unstack单级行索引:为何Store未转列索引,Product/Sales成列索引?
Pandas unstack 索引转换问题解答
问题重现
执行代码:
import pandas as pd data = { 'Store': ['Store_A', 'Store_A', 'Store_B', 'Store_B'], 'Product': ['Apples', 'Bananas', 'Apples', 'Bananas'], 'Sales': [200, 350, 150, 500] } df = pd.DataFrame(data) df.set_index(['Store'], inplace=True) df.unstack()
得到输出:
Store Product Store_A Apples Store_A Bananas Store_B Apples Store_B Bananas Sales Store_A 200 Store_A 350 Store_B 150 Store_B 500 dtype: object
疑问:为何作为行索引的Store未转换为列索引,反而Product和Sales成了列索引?预期结构为:以Product作为行索引,Store作为列索引,Sales为对应单元格的值(比如Apples行对应Store_A列值200、Store_B列值150)。
原因分析
unstack()的核心逻辑是将行索引中的指定层级转换为列索引。你先把Store设为唯一的行索引,此时行索引只有一层,调用unstack()时Pandas没有可转换的行索引层级,就会把原DataFrame的列(Product、Sales)当作新的行索引层级处理,最终出现和预期相反的结果。
正确实现方法
要得到预期结构,有两种简洁的实现方式:
方法1:set_index + unstack
先设置包含Store和Product的复合行索引,再将Store层级转成列索引:
import pandas as pd data = { 'Store': ['Store_A', 'Store_A', 'Store_B', 'Store_B'], 'Product': ['Apples', 'Bananas', 'Apples', 'Bananas'], 'Sales': [200, 350, 150, 500] } df = pd.DataFrame(data) # 设置Store和Product为复合行索引 df.set_index(['Store', 'Product'], inplace=True) # 将Store层级从行索引转成列索引 result = df.unstack(level='Store') print(result)
输出:
Sales Store Store_A Store_B Product Apples 200 150 Bananas 350 500
方法2:直接使用pivot
如果只需要最终的透视表结构,用pivot方法一步到位:
df.pivot(index='Product', columns='Store', values='Sales')
输出和上述结果完全一致。
内容的提问来源于stack exchange,提问作者mehran arbabian
相关产品推荐
相关产品推荐

