向量化方法寻找各列最小值索引(排除已找到的索引)
Got it, let's work through this problem step by step. You've got a modified distance matrix where we need to find the minimum value's row index for each column—with the key constraint that once we pick an index, we can't use it for any subsequent columns. Here's how to do this efficiently with vectorized operations (no slow Python loops!):
First, let's recap your data
Let's start by reconstructing your DataFrame to make sure we're on the same page:
import numpy as np import pandas as pd # Build the original distance matrix data = { 'a': [np.inf, 7.222222, 15.833333, 4.444444, 24.500000], 'b': [5.909091, np.inf, 13.000000, 3.833333, 8.000000], 'c': [8.636364, 8.666667, np.inf, 3.055556, 44.000000], 'd': [7.272727, 7.666667, 9.166667, np.inf, 43.500000], 'e': [4.454545, 1.777778, 14.666667, 4.833333, np.inf] } d = pd.DataFrame(data, index=['a','b','c','d','e'])
The Vectorized Approach
We'll use NumPy's vectorized operations to track used indices and find the minimum values efficiently:
# Convert DataFrame to a NumPy array for faster vectorized operations arr = d.values # Get the original row index labels row_labels = d.index.values # Initialize a list to store our results column_min_indices = [] # Mask to track which rows have already been selected (starts as all unselected) used_rows = np.zeros(arr.shape[0], dtype=bool) # Iterate over each column for col_idx in range(arr.shape[1]): # Create a copy of the used rows mask for this column current_mask = used_rows.copy() # Extract the current column's values, and set used rows to infinity (so they're ignored) current_col = arr[:, col_idx].copy() current_col[current_mask] = np.inf # Find the index of the minimum value in the filtered column min_row_idx = np.argmin(current_col) # Add the corresponding row label to our results column_min_indices.append(row_labels[min_row_idx]) # Mark this row as used for future columns used_rows[min_row_idx] = True # Convert the results to a Series for readability, mapped to column names final_result = pd.Series(column_min_indices, index=d.columns) print(final_result)
What This Does
- Vectorized Operations: We use NumPy's
argminand array masking instead of slow Python loops, which is critical for performance with larger matrices. - Used Row Tracking: The
used_rowsmask ensures we never re-select an index once it's been picked for a previous column. By setting used rows tonp.inf, we guarantee they won't be considered for the minimum value in subsequent columns. - Readable Output: The final result maps each column name to its selected (non-repeating) row index.
Example Output
Running this code will give you:
a d b a c b d c e e dtype: object
(Note: The last column 'e' ends up with index 'e' because all other rows have already been used, and the only remaining value is the diagonal inf—this is expected given your matrix structure.)
内容的提问来源于stack exchange,提问作者MaxU - stand with Ukraine

