如何解决Pandas中用DataFrame索引提取NumPy数组值的警告问题
我有一个NumPy数组:
import numpy as np np.random.seed(123456) a=np.random.randint(0,9, (5,5)) # 数组内容 # array([[1, 2, 1, 8, 0], # [7, 4, 8, 4, 2], # [6, 6, 7, 2, 6], # [2, 4, 4, 7, 4], # [4, 4, 5, 1, 7]])
还有一个存储目标数据索引的Pandas DataFrame:
import pandas as pd df = pd.DataFrame([[0,1],[1,2],[2,3]],columns=['i','j'])
过去用以下代码提取对应值:
df['vals'] = df[['i', 'j']].apply(lambda x: a[x[0], x[1]], axis=1)
能得到预期结果:
i j vals 0 0 1 2 1 1 2 8 2 2 3 2
但现在触发了FutureWarning:
FutureWarning: Series.getitem treating keys as positions is deprecated. In a future version, integer keys will always be treated as labels (consistent with DataFrame behavior). To access a value by position, use
ser.iloc[pos]
df['vals'] = df[['i', 'j']].apply(lambda x: a[x[0], x[1]], axis=1)
尝试过df.loc[:, 'vals'] = df.loc[:,['i', 'j']].apply(lambda x: a[x[0], x[1]], axis=1)仍有警告,用df['vals'] = df[['i', 'j']].apply(lambda x: a[*x], axis=1)出现语法错误,不想屏蔽警告,求解决方法。
方法1:用iloc明确按位置访问Series元素
警告的根源是x[0]、x[1]这种在Series中按位置取值的方式即将被弃用,改用iloc明确指定位置即可消除警告:
df['vals'] = df[['i', 'j']].apply(lambda x: a[x.iloc[0], x.iloc[1]], axis=1)
方法2:直接用NumPy数组索引(效率更高)
避开apply,直接提取i、j列的NumPy数组作为索引,大数据量下性能更优:
df['vals'] = a[df['i'].values, df['j'].values]
方法3:通过to_numpy转置后索引
将索引列转成二维数组后转置,再作为NumPy数组的索引:
df['vals'] = a[tuple(df[['i', 'j']].to_numpy().T)]
内容的提问来源于stack exchange,提问作者Wesley Kitlasten

