如何为numpy数组添加索引并创建可索引corr_matrix的ind_matrix
Hey there! Let's work through your problem step by step—sounds like you need a clean, indexable way to mark positions in your corr_matrix where enc_dict[enc_row] matches ret_dict[ret_col], plus some help with adding labels to numpy arrays. Here's how to do it:
1. Create ind_matrix as a Boolean Mask (No Slow Loops!)
Instead of iterating through rows and columns manually, we can use numpy's broadcasting to generate a boolean matrix that flags all matching positions in one go. This is way more efficient, especially for large matrices.
First, extract the values from your dictionaries that correspond to each row and column of corr_matrix:
import numpy as np # Assume your row indices map to enc_dict, column indices map to ret_dict row_values = np.array([enc_dict[row_idx] for row_idx in range(corr_matrix.shape[0])]) col_values = np.array([ret_dict[col_idx] for col_idx in range(corr_matrix.shape[1])])
Then generate the boolean mask using broadcasting (the [:, np.newaxis] turns the row array into a column vector so we can compare every row to every column):
ind_matrix = row_values[:, np.newaxis] == col_values
This ind_matrix will be the same shape as corr_matrix, with True wherever enc_dict[enc_row] == ret_dict[ret_col], and False otherwise.
2. Use ind_matrix to Index corr_matrix
Now you can use this mask directly to work with your correlation matrix:
- Extract all matching elements (returns a 1D array of the values):
matching_correlations = corr_matrix[ind_matrix] - Keep the original matrix shape but mask non-matching values (replace them with
NaNfor clarity):corr_matrix_filtered = corr_matrix.copy() corr_matrix_filtered[~ind_matrix] = np.nan # ~ flips True/False
3. Adding Index Labels to Your Numpy Array
If you want your array to have human-readable labels (like row/column names) instead of just integer indices, the easiest way is to convert it to a pandas DataFrame. This lets you work with labels directly, which simplifies filtering even more:
import pandas as pd # Convert corr_matrix to a DataFrame with labels from your dictionaries corr_df = pd.DataFrame( corr_matrix, index=[enc_dict[row_idx] for row_idx in range(corr_matrix.shape[0])], columns=[ret_dict[col_idx] for col_idx in range(corr_matrix.shape[1])] ) # Now you can filter rows where index matches column name in one line filtered_corr_df = corr_df.where(corr_df.index == corr_df.columns, np.nan)
If you absolutely need to stick with numpy, you could use a structured array, but pandas is far more intuitive for labeled data.
Quick Check Against Your Original Loop
If you were previously printing indices with nested loops like this:
for i in range(corr_matrix.shape[0]): for j in range(corr_matrix.shape[1]): if enc_dict[i] == ret_dict[j]: print(i, j)
The ind_matrix we created will exactly mark those (i,j) positions as True—so you're getting the same result, just in a format numpy can use directly for indexing.
内容的提问来源于stack exchange,提问作者Maria

