You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用RFECV结合对数变换目标时遇NaN错误的解决咨询

带对数变换目标的RFECV运行报错NaN问题解决

问题背景

用scikit-learn的RFECV做特征选择,目标是使用对y做对数变换的XGBoost模型(已验证效果更优)。未做目标变换的基础模型能正常运行RFECV,但用TransformedTargetRegressor实现对数变换后,运行时触发如下NaN错误:

ValueError: Input X contains NaN; RFECV does not accept missing values encoded as NaN natively. For supervised learning, you might want to consider sklearn.ensemble.HistGradientBoostingClassifier and Regressor which accept missing values encoded as NaNs natively. Alternatively, it is possible to preprocess the data, for instance by using an imputer transformer in a pipeline or drop samples with missing values. See https://scikit-learn.org/stable/modules/impute.html You can find a list of all estimators that handle NaN values at the following page: https://scikit-learn.org/stable/modules/impute.html#estimators-that-handle-nan-values

已确认y中无NaN,基础模型也不存在NaN问题,需解决该问题以实现带对数变换目标的RFECV运行。

问题原因

核心问题是TransformedTargetRegressor与RFECV的交互bug:RFECV需要通过estimator的feature_importances_属性获取特征重要性来做筛选,但TransformedTargetRegressor本身没有这个属性,虽然它会尝试转发请求给内部的XGBoost模型,但在部分scikit-learn版本中,这个转发过程会导致内部数据传递异常,进而误报X存在NaN。

此外需排除一种潜在情况:如果y中存在0或负数,np.log会生成-inf或NaN,但你已确认y无NaN,该情况可忽略。

解决方法

方案1:自定义包装类修复属性转发

写一个简单的包装类,继承TransformedTargetRegressor,显式实现feature_importances_属性,让RFECV能正确拿到内部XGBoost模型的特征重要性,这是最直接的解决方式。

补充验证步骤

如果仍报错,先做两个检查:

  • 确认X无NaN:执行print(np.isnan(x).any()),若返回True,先处理X的缺失值(比如用SimpleImputer填充)
  • 确认y全为正数:执行print((y <= 0).any()),若有0或负数,改用np.log1p和np.expm1做变换(避免log(0)产生-inf)

修改后的完整代码

import numpy as np
import xgboost as xgb
from sklearn.feature_selection import RFECV
from sklearn.compose import TransformedTargetRegressor
from scipy.stats import randint, uniform, loguniform

# 自定义包装类,确保feature_importances_正确转发给内部模型
class CustomTransformedTargetRegressor(TransformedTargetRegressor):
    @property
    def feature_importances_(self):
        return self.regressor_.feature_importances_

# 基础XGBoost模型
rs = 45
xgboost_reg = xgb.XGBRegressor(random_state = rs, 
                                grow_policy = "depthwise", 
                                booster = "gbtree",
                                tree_method = "auto",
                                n_estimators = randint(300,500).rvs(random_state = rs),
                                subsample = uniform(0.5, 0.5).rvs(random_state = rs),
                                max_depth = randint(3,10).rvs(random_state = rs),
                                learning_rate = loguniform(0.05, 0.2).rvs(random_state = rs),
                                colsample_bytree = uniform(0.5, 0.5).rvs(random_state = rs),
                                min_child_weight =  randint(1,20).rvs(random_state = rs),
                                gamma = uniform(0.5, 1).rvs(random_state = rs),
                                reg_alpha = uniform(0.0, 1.0).rvs(random_state = rs),
                                reg_lambda = uniform(0.0, 1.0).rvs(random_state = rs),
                                max_delta_step = randint(1,10).rvs(random_state = rs)
)

# RFECV配置
step = 20
min_features_to_select = 9

# 使用自定义包装类创建带对数变换的模型
log_estimator = CustomTransformedTargetRegressor(regressor=xgboost_reg,
                                                 func=np.log,
                                                 inverse_func=np.exp)
# 运行RFECV
rfecv_log = RFECV(
    estimator=log_estimator,
    step=step,
    cv=4,
    scoring="neg_root_mean_squared_error",
    min_features_to_select=min_features_to_select,
    n_jobs=-1, 
)
rfecv_log.fit(x, y)
print(rfecv_log.n_features_)

内容的提问来源于stack exchange,提问作者hexolitemax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 07:40:55