You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Sklearn回归模型处理连续浮点标签报Unknown label type问题分析

问题:回归模型Pipeline中触发Unknown label type错误

在使用Scikit-learn的RandomForestRegressor等回归模型结合Pipeline处理连续浮点型标签时,出现了原本只在分类模型处理实值标签时才会触发的ValueError: Unknown label type错误。

标签是经过log变换的ICU住院天数(为避免0值产生负无穷,已添加小偏移量),原始整数标签可正常训练,但变换后的浮点标签报错。排查过程:

  • 将标签转为整数后模型正常运行,排除模型本身问题
  • 仅使用Pipeline时才报错,问题定位到预处理流程
  • 移除分类特征预处理中的SelectPercentile(chi2)步骤后,模型可正常拟合,但不清楚原因

训练代码

train_X = df_train_trimmed.loc[:, ~df_train_trimmed.columns.isin(['log_icu_days'])]
train_y = list(df_train_trimmed["log_icu_days"].astype(np.float32))
rfr = Pipeline(
    [
        ("preprocessor", regressor_preprocessor),
        (
            "regressor",
            RandomForestRegressor(n_estimators=1200, max_depth=30, criterion="squared_error", random_state=0, n_jobs=-1)
        ),
    ]
)
rfr.fit(
    train_X, train_y
)

报错信息

C:\ProgramData\Anaconda3\lib\site-packages\sklearn\utils\multiclass.py in unique_labels(*ys)
    105     _unique_labels = _FN_UNIQUE_LABELS.get(label_type, None)
    106     if not _unique_labels:
--> 107         raise ValueError("Unknown label type: %s" % repr(ys))
    108 
    109     if is_array_api:

ValueError: Unknown label type: (array([ 2.19833517, -4.60517025, -4.60517025, ..., -4.60517025, -4.60517025, -4.60517025]),)

预处理Pipeline代码

categorical_features = [
    "sex", "approach", "preop_pft"
]
categorical_transformer = Pipeline(
    steps=[
        ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),
        ("selector", SelectPercentile(chi2, percentile=50)),
    ]
)

numeric_features = list(set(df_train_trimmed.columns) - set(categorical_features) - {'icu_days', 'log_icu_days'})
numeric_transformer = Pipeline(
    steps=[("imputer", SimpleImputer(strategy="median")), ("scaler", StandardScaler())]
)

regressor_preprocessor = ColumnTransformer(
    transformers=[
        ("num", numeric_transformer, numeric_features),
        ("cat", categorical_transformer, categorical_features),
    ],
    remainder="passthrough"
)

原因分析

SelectPercentile(chi2)是分类任务专用的特征选择方法,它依赖卡方检验来衡量特征和离散分类标签的相关性。当你在回归任务的预处理流程中使用它时,Scikit-learn内部会错误地将连续浮点型的回归标签当作分类标签处理,而卡方检验无法适配连续标签,最终触发Unknown label type错误。

解决方案

方法1:替换为回归专用的特征选择器

把SelectPercentile(chi2)换成适合回归任务的SelectPercentile(f_regression),它专门用于衡量特征与连续标签的线性相关性,适配回归场景:

from sklearn.feature_selection import SelectPercentile, f_regression

categorical_transformer = Pipeline(
    steps=[
        ("encoder", OneHotEncoder(handle_unknown="ignore", sparse_output=False)),
        ("selector", SelectPercentile(f_regression, percentile=50)),
    ]
)

方法2:移除特征选择步骤

如果不需要对分类特征进行筛选,直接去掉SelectPercentile步骤即可,这也是你之前测试验证有效的方法。


内容的提问来源于stack exchange,提问作者Ryan Folks

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:00:35