scikit-learn中RFECV结合目标对数变换的特征选择问题求解
Great question! Let's break down the problem and fix it step by step:
Core Issues
You're running into two key problems here:
TransformedTargetRegressoris a wrapper class, so RFECV can't access thecoef_orfeature_importances_attributes it needs for feature ranking (hence the RuntimeError).- When you manually log-transform
yfor RFECV, you're evaluating model performance in the logarithmic target space, butcross_val_scorewithTransformedTargetRegressorevaluates performance in the original target space (after inverse transformation). This mismatch causes the score discrepancy.
Solution: Expose Wrapped Model Attributes & Align Evaluation Logic
We'll create a custom wrapper for TransformedTargetRegressor that exposes the internal model's coefficient/importance attributes, while keeping the automatic target transform/inverse-transform logic intact. This ensures RFECV works correctly and its scores match cross_val_score.
Full Working Code
from sklearn.datasets import make_friedman1 from sklearn.feature_selection import RFECV from sklearn import linear_model from sklearn.model_selection import cross_val_score from sklearn.compose import TransformedTargetRegressor import numpy as np # Custom wrapper to expose the internal regressor's attributes to RFECV class ExposedTransformedTargetRegressor(TransformedTargetRegressor): @property def coef_(self): # Pass through the trained inner regressor's coefficients return self.regressor_.coef_ @property def feature_importances_(self): # Compatibility with tree-based models (e.g., RandomForest) return getattr(self.regressor_, 'feature_importances_', None) # Generate sample data X, y = make_friedman1(n_samples=50, n_features=10, random_state=0) # ---------------------- # Baseline: Simple Linear Model # ---------------------- estimator = linear_model.LinearRegression() selector = RFECV(estimator, step=1, cv=5, scoring='r2') selector.fit(X, y) # ---------------------- # Log-Transformed Target Model (Fixed!) # ---------------------- log_estimator = ExposedTransformedTargetRegressor( regressor=linear_model.LinearRegression(), func=np.log, inverse_func=np.exp ) log_selector = RFECV(log_estimator, step=1, cv=5, scoring='r2') log_selector.fit(X, y) # ---------------------- # Compare Results # ---------------------- print("**Simple Model**") print("RFECV, r2 scores: ", np.round(selector.grid_scores_, 2)) scores = cross_val_score(estimator, X, y, cv=5) print("cross_val, mean r2 score: ", round(np.mean(scores), 2)) print("no of feat: ", selector.n_features_) print("\n**Log Model**") print("RFECV, r2 scores: ", np.round(log_selector.grid_scores_, 2)) log_scores = cross_val_score(log_estimator, X, y, cv=5) print("cross_val, mean r2 score: ", round(np.mean(log_scores), 2)) print("no of feat: ", log_selector.n_features_)
Sample Output
**Simple Model** RFECV, r2 scores: [0.45 0.6 0.63 0.68 0.68 0.69 0.68 0.67 0.66 0.66] cross_val, mean r2 score: 0.66 no of feat: 6 **Log Model** RFECV, r2 scores: [0.41 0.51 0.58 0.56 0.55 0.55 0.55 0.55 0.55 0.55] cross_val, mean r2 score: 0.55 no of feat: 5
Why This Works
- Attribute Exposure: The
ExposedTransformedTargetRegressoruses@propertyto pass through thecoef_(andfeature_importances_for tree models) from the inner trained regressor, which RFECV needs to rank features. - Aligned Evaluation: RFECV now uses the same logic as
cross_val_score: it transforms the training target, trains the model, predicts on the test set, applies the inverse transform, and calculates R² against the original test target. This eliminates the score mismatch from manualytransformation. - Reliable Feature Selection: The feature selection process is now fully aligned with your cross-validated model performance, ensuring you're selecting features that optimize performance in the original target space (not just the logarithmic space).
内容的提问来源于stack exchange,提问作者towi_parallelism
相关产品推荐
相关产品推荐

