如何按行选择Pandas DataFrame的不同列以提取矩阵?
从Pandas DataFrame按指定列矩阵提取对应值
给定如下Pandas DataFrame:
a b c 0 1 2 3 1 4 5 6 2 7 8 9 3 10 11 12 4 13 14 15
需要根据以下提取矩阵(每行指定要提取的列名):
[ ['a', 'a'], ['a', 'a'], ['a', 'b'], ['c', 'a'], ['b', 'b'] ]
提取得到对应的数值矩阵:
[ [1, 1], [4, 4], [7, 8], [12, 10], [14, 14] ]
由于原DataFrame规模极大,pd.iterrows()速度太慢,需基于充足内存实现高效处理。
高效实现方案
方法1:列名转索引 + NumPy高级索引(推荐)
通过将列名映射为整数索引,结合NumPy的高级索引直接提取,全程向量化操作,无循环:
import pandas as pd import numpy as np # 构造原DataFrame df = pd.DataFrame({ 'a': [1, 4, 7, 10, 13], 'b': [2, 5, 8, 11, 14], 'c': [3, 6, 9, 12, 15] }) # 定义提取矩阵 col_matrix = np.array([ ['a', 'a'], ['a', 'a'], ['a', 'b'], ['c', 'a'], ['b', 'b'] ]) # 建立列名到索引的映射字典 col_to_idx = {col: idx for idx, col in enumerate(df.columns)} # 将提取矩阵的列名转为对应整数索引 idx_matrix = np.vectorize(col_to_idx.get)(col_matrix) # 生成行索引矩阵(每行对应原DataFrame的行号) row_indices = np.arange(df.shape[0])[:, np.newaxis] # 用NumPy索引直接提取值 result = df.to_numpy()[row_indices, idx_matrix] print(result)
输出结果:
[[ 1 1] [ 4 4] [ 7 8] [12 10] [14 14]]
方法2:使用DataFrame.lookup(旧版Pandas兼容)
若使用Pandas 1.2.x及更早版本,可直接用lookup方法,需将提取矩阵拆分为两个列序列:
# 拆分提取矩阵为两个Series col_left = pd.Series(col_matrix[:, 0]) col_right = pd.Series(col_matrix[:, 1]) # 分别提取两列对应的值 vals_left = df.lookup(df.index, col_left) vals_right = df.lookup(df.index, col_right) # 组合成结果矩阵 result_matrix = np.column_stack([vals_left, vals_right]) print(result_matrix)
性能说明
两种方法均为向量化操作,完全规避了iterrows这类低效循环,在大规模数据场景下速度提升显著。其中方法1依托NumPy的数组索引机制,内存利用更高效,是优先推荐的方案。
内容的提问来源于stack exchange,提问作者John Rouhana
相关产品推荐
相关产品推荐

