如何按每行列索引列表选取pandas中data2的对应元素?
How to Select Elements from data2 Using Row-Wise idxmax from data1
Got it, let's tackle this problem—you want to extract values from data2 using the column indices where data1 has its row-wise maximum. Here are a few practical, efficient methods to get your desired result:
Method 1: Use numpy.take_along_axis (Best for Large Datasets)
This is the fastest approach because it uses vectorized NumPy operations, which avoid slow row-by-row iteration. Perfect if you're working with big data.
import pandas as pd import numpy as np data1 = pd.DataFrame([[1,2], [4,3], [5,6]]) data2 = pd.DataFrame([[10,20], [30,40], [50,60]]) # Get the column indices of row-wise maxima from data1 max_col_indices = data1.idxmax(axis=1).values.reshape(-1, 1) # Reshape to 2D so it aligns with data2's structure for take_along_axis # Extract the corresponding values from data2 selected_values = np.take_along_axis(data2.values, max_col_indices, axis=1).flatten() # Convert to a pandas Series (or DataFrame if needed) result = pd.Series(selected_values, name="selected_values") print(result)
Output:
0 20 1 30 2 60 Name: selected_values, dtype: int64
Method 2: Use DataFrame.lookup (Concise for Small Data)
This method is super straightforward, though note it's marked as deprecated in pandas 1.2+. Still works great for smaller datasets:
max_cols = data1.idxmax(axis=1) # Lookup values using data2's row indices and the max column indices from data1 result = pd.Series(data2.lookup(data2.index, max_cols), name="selected_values")
Method 3: Use apply (Intuitive but Slow for Big Data)
If you're dealing with a tiny dataset and prioritize readability over speed, this row-wise approach works—but avoid it for large data:
max_cols = data1.idxmax(axis=1) result = data2.apply(lambda row: row[max_cols[row.name]], axis=1)
Quick Breakdown:
- All methods start with
data1.idxmax(axis=1)to get the column where each row indata1has its maximum value. take_along_axisaligns the index array withdata2's underlying values to pull the right elements efficiently.lookupdirectly maps row and column positions, but keep an eye on its deprecation status for future pandas versions.applyiterates over each row, which is easy to read but much slower for large datasets.
内容的提问来源于stack exchange,提问作者Kibeom Kim
相关产品推荐
相关产品推荐

