如何批量计算Pandas数据集互信息?多因变量场景解决方案
Got it, let's fix that tedious manual repetition and automate both your mutual information tasks efficiently!
First, let's handle computing normalized mutual information between each of your 46 dependent variables and the 4 independent features. Here's a clean, automated approach:
Step-by-Step Code:
import pandas as pd import numpy as np from sklearn.feature_selection import mutual_info_regression # Define your independent features (X) X = df[['Temperature', 'Precipitation', 'Dew', 'Snow']] # Get all dependent variables (all columns not in X) dependent_cols = [col for col in df.columns if col not in X.columns] # Create an empty DataFrame to store results (rows = features, cols = dependent vars) mi_results = pd.DataFrame(index=X.columns) # Loop through each dependent variable and compute mutual info for y_col in dependent_cols: # Extract the target as a 1D array (mutual_info_regression expects 1D y) y = df[y_col].squeeze() # Calculate mutual info scores mi_scores = mutual_info_regression(X, y) # Normalize scores by the maximum value (matching your original code) mi_normalized = mi_scores / np.max(mi_scores) # Add the normalized scores to our results DataFrame mi_results[y_col] = mi_normalized # Optional: Sort results for easier reading (e.g., sort features by a specific dependent var) # mi_results_sorted = mi_results.sort_values(by='N0037', ascending=False) print(mi_results)
This code will churn through all 46 dependent variables in one run, no manual variable swapping needed. The result is a DataFrame where each column corresponds to a dependent variable, and each row shows the normalized mutual info with your 4 features.
If you want to compute mutual information between every pair of variables (including between dependent variables, and between independent features), we can build a symmetric pairwise matrix. Here's how:
Step-by-Step Code:
# Get all columns in your dataset all_cols = df.columns # Initialize an empty square matrix to store pairwise mutual info pairwise_mi = pd.DataFrame(index=all_cols, columns=all_cols, dtype=np.float64) # Fill the matrix (mutual info is symmetric, so we compute each pair once) for i in range(len(all_cols)): col_a = all_cols[i] for j in range(i, len(all_cols)): col_b = all_cols[j] if col_a == col_b: # Mutual info of a variable with itself is its entropy, but we'll set to 0 for simplicity # If you want entropy instead, use sklearn.metrics.entropy (requires discretization) pairwise_mi.loc[col_a, col_b] = 0.0 else: # Compute mutual info between col_a and col_b # Treat one as feature, the other as target mi_score = mutual_info_regression(df[[col_a]], df[col_b].squeeze())[0] # Assign to both positions since I(A,B) = I(B,A) pairwise_mi.loc[col_a, col_b] = mi_score pairwise_mi.loc[col_b, col_a] = mi_score # Optional: Normalize the entire matrix if you want scores between 0 and 1 # pairwise_mi_normalized = pairwise_mi / pairwise_mi.max().max() print(pairwise_mi)
Key Notes:
- Symmetry: Mutual information is symmetric, so we only compute each pair once to save computation time.
- Diagonal Elements: By default, we set these to 0, but if you want the entropy of each variable (which is the mutual info of a variable with itself), you can use
sklearn.metrics.entropy—just note that continuous variables need to be discretized first (similar to howmutual_info_regressionworks under the hood). - Continuous Variables:
mutual_info_regressionuses k-nearest neighbors to estimate mutual info for continuous data, which is perfect for your dataset.
内容的提问来源于stack exchange,提问作者Denise

