You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn pipeline中将PCA结果与原始特征拼接的更简洁实现方案

更简洁的实现方案

方案1:使用FunctionTransformer实现恒等变换(最直观)

你需要的逻辑是原样输出所有输入特征,直接用scikit-learn内置的FunctionTransformer即可,它默认就是不做任何修改返回输入数据,完全可以替代你写的ColumnTransformer,代码修改后如下:

from sklearn.preprocessing import FunctionTransformer
from sklearn.pipeline import FeatureUnion, make_pipeline
from sklearn.decomposition import PCA

feature_selector = FeatureUnion(
    [
        ("original", FunctionTransformer()),  # 原样返回所有输入特征
        ("pca", PCA())
    ])
my_pipeline = make_pipeline(preprocessor, feature_selector, model)

这个方案完全符合scikit-learn的API规范,支持pipeline的fit、transform全流程,也兼容特征名追溯的get_feature_names_out方法。

方案2:使用全选特征选择器SelectKBest

如果你希望用特征选择器类的组件实现全列选择,可以用SelectKBest指定k="all",它会直接返回所有输入特征,也能满足需求:

from sklearn.feature_selection import SelectKBest

feature_selector = FeatureUnion(
    [
        ("original", SelectKBest(k="all")),  # 选择全部特征输出
        ("pca", PCA())
    ])

补充注意事项

PCA对特征尺度高度敏感,如果你的preprocessor步骤没有做标准化处理,建议给PCA部分加上标准化步骤,避免特征尺度差异影响PCA效果:

("pca", make_pipeline(StandardScaler(), PCA()))

内容的提问来源于stack exchange,提问作者abdelgha4

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 06:24:04