使用sklearn SFS预测时出现'Feature names seen at fit time, yet now missing'错误
问题描述
使用sklearn的Sequential Feature Selection(SFS)选中特征子集后,对X_test进行预测时遇到错误:
Feature names seen at fit time, yet now missing
相关代码如下:
model_for_sfs = LogisticRegression(solver="saga") model = LogisticRegression(solver="saga") pipeline_for_fs = Pipeline(steps=[ ('imputer', SimpleImputer(strategy="median")), ("model",model_for_sfs)]) n_splits = 2 cv_fs = StratifiedKFold(n_splits, shuffle=True, random_state=0) cv_perf = StratifiedKFold(n_splits, shuffle=True, random_state=0) # Feature selection fs = SFS( estimator=pipeline_for_fs, n_features_to_select=2, cv=cv_fs, scoring='accuracy', n_jobs=-1 ) pipeline = Pipeline(steps=[ ('imputer', SimpleImputer(strategy="median")), ('selector', fs), ("model", model)]) pipeline.fit(X_train, y_train) sfs = pipeline.named_steps["selector"] features = sfs.get_support(indices=True) y_pred = pipeline.predict_proba(X_test.iloc[:, list(features)])[:, 1]
原本以为sklearn SFS会自动转换数据集以仅保留所选特征,请问实际情况是否如此?有没有方法实现这一功能?
解决方案
关于SFS自动转换数据集的实际情况
你的理解存在偏差:SFS本身会在内部处理特征选择,但当前代码的Pipeline逻辑出现了冲突:
- 你构建的完整Pipeline已经包含了
缺失值填充→特征选择→模型预测的全流程,pipeline.fit()后,整个Pipeline已经学习了从原始特征到最终预测的所有转换规则。 - 但你最后手动对
X_test做切片筛选特征,跳过了Pipeline中imputer和selector的预处理逻辑,直接给模型喂了未经过统一预处理的特征子集,导致特征维度/名称与模型训练时不匹配,触发错误。
正确实现自动特征转换的方法
直接调用训练好的pipeline的predict_proba方法即可,无需手动处理X_test:
y_pred = pipeline.predict_proba(X_test)[:, 1]
整个Pipeline会自动执行以下步骤:
- 用训练阶段拟合好的
SimpleImputer对X_test做缺失值填充 - 用训练好的
SFS自动筛选出选中的特征子集 - 用最终的
LogisticRegression模型生成预测概率
额外说明
如果需要查看选中的特征信息,可以使用get_support,但不要用它手动处理测试集:
# 获取选中的特征索引(仅用于查看) selected_indices = pipeline.named_steps["selector"].get_support(indices=True) # 若X_train是DataFrame,可获取特征名称 selected_names = X_train.columns[selected_indices]
内容的提问来源于stack exchange,提问作者Charlie
相关产品推荐
相关产品推荐

