如何按索引保留Pandas DataFrame中对应列的数值?
高效实现Pandas索引-列映射提取数据
原始DataFrame
col_a col_b col_c 0 10 15 20 0 10 15 20 1 10 15 20 1 10 15 20 1 10 15 20 1 10 15 20 2 10 15 20
索引-列映射规则
{ 0: 'col_a', 1: 'col_b', 2: 'col_c' }
目标输出
column 0 10 0 10 1 15 1 15 1 15 1 15 2 20
现有实现(性能待优化)
keep_cols = [(0, 'col_a'), (1, 'col_b'), (2, 'col_c')] output = pd.concat([df.loc[df['index_col'] == idx, col] for idx, col in keep_cols], axis=1)
优化方案
方案1:使用df.lookup(高效且直接)
lookup是Pandas专门用于按行索引+列名匹配提取值的方法,完全避免循环拼接,性能最优:
map_dict = {0: 'col_a', 1: 'col_b', 2: 'col_c'} # 为每行匹配对应的目标列名 target_cols = df.index.map(map_dict) # 提取对应值并转为目标格式 output = pd.DataFrame(df.lookup(df.index, target_cols), index=df.index, columns=['column'])
方案2:numpy向量化批量赋值
利用numpy的布尔索引实现批量操作,同样避免循环切片,性能接近lookup:
import numpy as np map_dict = {0: 'col_a', 1: 'col_b', 2: 'col_c'} result = np.zeros(len(df), dtype=int) # 遍历映射规则,批量赋值 for idx, col in map_dict.items(): mask = df.index == idx result[mask] = df.loc[mask, col].values output = pd.DataFrame(result, index=df.index, columns=['column'])
方案3:apply简洁实现(代码短但性能略逊)
如果优先追求代码简洁,可使用apply,但数据量大时性能不如前两种向量化方法:
map_dict = {0: 'col_a', 1: 'col_b', 2: 'col_c'} output = df.apply(lambda row: row[map_dict[row.name]], axis=1).to_frame('column')
以上方案均避免了多次子DataFrame切片与拼接操作,在数据量较大时性能提升明显,其中lookup和numpy向量化方案的执行效率最高。
内容的提问来源于stack exchange,提问作者theodosis
相关产品推荐
相关产品推荐

