Python中如何按市政区分组查看随机森林回归模型的特征重要性以分析空间差异
Got it, let's break down how you can solve this problem—whether you want to train separate models per municipality or use your existing global model to uncover per-municipality feature importance patterns. Here are three actionable, hands-on approaches:
This is the most straightforward method: split your data by municipality, train a mini random forest for each group, then collect feature importance for every area. It works best if you have enough samples per municipality to train a reliable model.
Steps:
- Group your dataset by the
municipalitycolumn. - For each group, train a RandomForestRegressor (reuse your existing hyperparameters to keep consistency!).
- Extract the feature importance array from each trained model and map it to the corresponding municipality.
- Aggregate results into a clean DataFrame for analysis.
Code Example:
import pandas as pd from sklearn.ensemble import RandomForestRegressor # Assume your data is stored in `df`, with features in `X_cols` and target in `y_col` X_cols = [col for col in df.columns if col not in ['municipality', 'energy_consumption']] y_col = 'energy_consumption' # Initialize a dictionary to store results municipality_feature_importance = {} # Group data by municipality and train models for mun, group in df.groupby('municipality'): # Skip small groups if needed (adjust threshold based on your data) if len(group) < 50: continue X = group[X_cols] y = group[y_col] # Train model with your existing hyperparameters rf = RandomForestRegressor(n_estimators=100, random_state=42) rf.fit(X, y) # Store feature importance as a Series for easy alignment municipality_feature_importance[mun] = pd.Series(rf.feature_importances_, index=X_cols) # Convert to a DataFrame for quick analysis mun_importance_df = pd.DataFrame(municipality_feature_importance).T # Check top 3 features for a specific municipality print(mun_importance_df.loc['Amsterdam'].sort_values(ascending=False).head(3))
If you don't want to train separate models (e.g., some municipalities have too few samples), SHAP is a powerful tool. It calculates how much each feature contributes to the prediction of every single sample using your pre-trained global random forest. You can then average these contributions per municipality to get a representative feature importance for each area.
Steps:
- Load your pre-trained RandomForestRegressor.
- Use SHAP's
TreeExplainerto compute SHAP values for all samples. - Merge the SHAP values with your original data (including the
municipalitycolumn). - Aggregate absolute SHAP values per municipality (since feature importance focuses on magnitude, not direction of impact).
Code Example:
import shap import pandas as pd # Assume your pre-trained global model is `global_rf` and full dataset is `df` X = df[X_cols] explainer = shap.TreeExplainer(global_rf) shap_values = explainer.shap_values(X) # Convert SHAP values to a DataFrame shap_df = pd.DataFrame(shap_values, columns=X_cols) # Add municipality column to link results to areas shap_df['municipality'] = df['municipality'] # Calculate mean absolute SHAP value per municipality (your per-municipality feature importance) mun_shap_importance = shap_df.groupby('municipality').apply(lambda x: x[X_cols].abs().mean()) # Check top 3 features for a specific municipality print(mun_shap_importance.loc['Rotterdam'].sort_values(ascending=False).head(3))
If you specifically want to find which feature is the most impactful for the majority of samples in each municipality, you can modify the SHAP approach to pick the top feature per sample, then count occurrences per area.
Code Example:
# For each sample, find the feature with the largest absolute SHAP value shap_df['top_feature'] = shap_df[X_cols].abs().idxmax(axis=1) # Count how many times each feature is the top one per municipality (normalized to percentages) top_feature_counts = shap_df.groupby('municipality')['top_feature'].value_counts(normalize=True) # Get the most common top feature for each municipality most_common_top_feature = top_feature_counts.groupby('municipality').head(1) print(most_common_top_feature)
- Approach 1 is ideal if you have sufficient samples per municipality (aim for 50-100+ samples) and want model-specific importance tailored to each area.
- Approach 2 works best for uneven sample sizes—you get the generalization power of your global model while still capturing per-municipality patterns.
- Bonus Approach directly implements your original idea, giving you a count-based view of which features drive predictions most often in each area.
内容的提问来源于stack exchange,提问作者Brandon

