如何在Pandas DataFrame中按行选取绝对值最大且保留符号的值
问题:为数值型DataFrame每行选取绝对值最大且保留符号的值
我有一个仅包含数值数据的DataFrame:
[ In1]: df = pd.DataFrame(np.random.randn(5, 3).round(2), columns=['A', 'B', 'C']) df [Out1]: A B C 0 -0.27 1.22 1.10 1 -3.22 0.48 -1.64 2 1.42 0.24 -0.12 3 -1.12 0.44 0.23 4 1.88 -0.38 0.62
需要为每行选取绝对值最大且保留原始符号的值,预期结果如下:
0 1.22 1 -3.22 2 1.42 3 -1.12 4 1.88
我已经通过以下代码确定了每行最大值所在的列:
[ In2]: loc_max = df.abs().idxmax(axis=1) loc_max [Out2]: 0 B 1 A 2 A 3 A 4 A
由于实际使用的DataFrame数据量极大,性能是核心需求。
解决方案与性能对比
以下四种方法均能得到预期结果,我们在规模更大的DataFrame(1000行×100列)上进行性能测试:
测试代码
df = pd.DataFrame(np.random.randn(1000, 100).round(2)) def numpy_argmax(): idx_max = np.abs(df.values).argmax(axis=1) val = df.values[range(len(df)), idx_max] return pd.Series(val, index=df.index) def check_sign(): row_max = df.abs().max(axis=1) return row_max * (-1) ** df.ne(row_max, axis=0).all(axis=1) def loop_rows(): return df.apply(lambda s: s[s.abs().idxmax()], axis=1) def pandas_loc(): s = df.abs().idxmax(axis=1) val = [df.loc[x, y] for x, y in zip(s.index, s)] return pd.Series(val, index=df.index) %timeit numpy_argmax() %timeit check_sign() %timeit loop_rows() %timeit pandas_loc()
性能测试结果(从快到慢排序)
numpy_argmax():性能最优,直接调用底层numpy实现,避免Pandas的额外开销check_sign():次之,通过计算符号位得到结果pandas_loc():性能一般,通过遍历索引取值loop_rows():性能最差,使用apply逐行处理,效率极低
结论
直接基于Pandas底层的numpy实现是处理大规模数据时的最优选择,能最大化性能。
内容的提问来源于stack exchange,提问作者data-monkey
相关产品推荐
相关产品推荐

