Pandas 2.1中stack()与unstack()函数报错问题求助
Pandas 2.1.0中MultiIndex列/索引执行
stack(future_stack=True)报错问题及解决 问题现象
现有数据框df:
(O0, O, 0.66) (H1, H, 0.31) (H2, H, 0.31) (O0, O, 0.66) NaN NaN NaN (H1, H, 0.31) 0.957897 NaN NaN (H2, H, 0.31) 0.957897 1.5109 NaN
执行df.stack(future_stack=True)时触发报错:
KeyError: "None of [Index(['O0'], dtype='object')] are in the [columns]"
调用df.unstack()也会出现相同报错。该操作在Pandas 1.5.3中可正常运行,Pandas 2.1.0版本失效,问题源于索引和列的格式处理变更。
数据框生成逻辑
该数据框由以下代码生成:
df=pd.DataFrame(squareform(pdist(xyz_df.iloc[:,3:6],'euclid')), columns = xyz_df[['atom1_idx','element','cov_radius']], index = xyz_df[['atom2_idx','element','cov_radius']])
xyz_df.iloc[:,3:6]存储原子坐标squareform(来自scipy)用于将距离向量转换为对称距离矩阵
临时可行方案
使用df.stack(future_stack=False).dropna()可正常执行,输出结果如下:
(H1, H, 0.31) (O0, O, 0.66) 0.957897 (H2, H, 0.31) (O0, O, 0.66) 0.957897 (H1, H, 0.31) 1.510900
最终期望得到的格式:
level_1 0 (H1, H, 0.31) (O0, O, 0.66) 0.957897 (H2, H, 0.31) (O0, O, 0.66) 0.957897 (H2, H, 0.31) (H1, H, 0.31) 1.510900
报错复现代码
以下极简代码可触发相同报错:
import pandas as pd import numpy as np c_1 = 'A0', 'B0', 0.11 c_2 = 'A1', 'B1', 0.12 c_3 = 'A2', 'B2', 0.14 df=pd.DataFrame(np.random.rand(3,3), columns= [c_1,c_2,c_3], index = ['(i1, i2, i3)','(i4, i5, i6)','(i7, i8, i9)']) print(df) # 正常执行 print(df.stack(future_stack=False)) print(df.unstack()) # 触发报错 print(df.stack(future_stack=True))
官方修复进展
经确认这是Pandas的已知bug,预计在2.1.1版本中完成修复。
内容的提问来源于stack exchange,提问作者rsdel
相关产品推荐
相关产品推荐

