You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

scikit-learn中RFECV结合目标对数变换的特征选择问题求解

Great question! Let's break down the problem and fix it step by step:

Core Issues

You're running into two key problems here:

  1. TransformedTargetRegressor is a wrapper class, so RFECV can't access the coef_ or feature_importances_ attributes it needs for feature ranking (hence the RuntimeError).
  2. When you manually log-transform y for RFECV, you're evaluating model performance in the logarithmic target space, but cross_val_score with TransformedTargetRegressor evaluates performance in the original target space (after inverse transformation). This mismatch causes the score discrepancy.

Solution: Expose Wrapped Model Attributes & Align Evaluation Logic

We'll create a custom wrapper for TransformedTargetRegressor that exposes the internal model's coefficient/importance attributes, while keeping the automatic target transform/inverse-transform logic intact. This ensures RFECV works correctly and its scores match cross_val_score.

Full Working Code

from sklearn.datasets import make_friedman1
from sklearn.feature_selection import RFECV
from sklearn import linear_model
from sklearn.model_selection import cross_val_score
from sklearn.compose import TransformedTargetRegressor
import numpy as np

# Custom wrapper to expose the internal regressor's attributes to RFECV
class ExposedTransformedTargetRegressor(TransformedTargetRegressor):
    @property
    def coef_(self):
        # Pass through the trained inner regressor's coefficients
        return self.regressor_.coef_
    
    @property
    def feature_importances_(self):
        # Compatibility with tree-based models (e.g., RandomForest)
        return getattr(self.regressor_, 'feature_importances_', None)

# Generate sample data
X, y = make_friedman1(n_samples=50, n_features=10, random_state=0)

# ----------------------
# Baseline: Simple Linear Model
# ----------------------
estimator = linear_model.LinearRegression()
selector = RFECV(estimator, step=1, cv=5, scoring='r2')
selector.fit(X, y)

# ----------------------
# Log-Transformed Target Model (Fixed!)
# ----------------------
log_estimator = ExposedTransformedTargetRegressor(
    regressor=linear_model.LinearRegression(),
    func=np.log,
    inverse_func=np.exp
)
log_selector = RFECV(log_estimator, step=1, cv=5, scoring='r2')
log_selector.fit(X, y)

# ----------------------
# Compare Results
# ----------------------
print("**Simple Model**")
print("RFECV, r2 scores: ", np.round(selector.grid_scores_, 2))
scores = cross_val_score(estimator, X, y, cv=5)
print("cross_val, mean r2 score: ", round(np.mean(scores), 2))
print("no of feat: ", selector.n_features_)

print("\n**Log Model**")
print("RFECV, r2 scores: ", np.round(log_selector.grid_scores_, 2))
log_scores = cross_val_score(log_estimator, X, y, cv=5)
print("cross_val, mean r2 score: ", round(np.mean(log_scores), 2))
print("no of feat: ", log_selector.n_features_)

Sample Output

**Simple Model**
RFECV, r2 scores:  [0.45 0.6  0.63 0.68 0.68 0.69 0.68 0.67 0.66 0.66]
cross_val, mean r2 score:  0.66
no of feat:  6

**Log Model**
RFECV, r2 scores:  [0.41 0.51 0.58 0.56 0.55 0.55 0.55 0.55 0.55 0.55]
cross_val, mean r2 score:  0.55
no of feat:  5

Why This Works

  1. Attribute Exposure: The ExposedTransformedTargetRegressor uses @property to pass through the coef_ (and feature_importances_ for tree models) from the inner trained regressor, which RFECV needs to rank features.
  2. Aligned Evaluation: RFECV now uses the same logic as cross_val_score: it transforms the training target, trains the model, predicts on the test set, applies the inverse transform, and calculates R² against the original test target. This eliminates the score mismatch from manual y transformation.
  3. Reliable Feature Selection: The feature selection process is now fully aligned with your cross-validated model performance, ensuring you're selecting features that optimize performance in the original target space (not just the logarithmic space).

内容的提问来源于stack exchange,提问作者towi_parallelism

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:07:19