You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从包含独热编码与MinMax缩放的ColumnTransformer实例中获取特征名称列表

获取ColumnTransformer处理后的所有特征名称

Got it, let's walk through how to get the full list of feature names after your ColumnTransformer finishes applying OneHotEncoder and MinMaxScaler. There are two reliable approaches depending on your scikit-learn version:

方法1:使用scikit-learn 0.23+(推荐)

Starting from scikit-learn 0.23, ColumnTransformer has a built-in get_feature_names_out() method that handles all the heavy lifting for you—no manual stitching needed.

Here's how to use it:

# 首先确保你的ColumnTransformer已经拟合过原始数据
ct.fit(X)  # X是你的输入数据集

# 获取所有处理后的特征名称
processed_feature_names = ct.get_feature_names_out()

# 查看结果
print(processed_feature_names)

结果说明:

  • OneHotEncoder生成的特征会带上你定义的前缀onehot__,格式为onehot__原始类别特征名_类别值
  • MinMaxScaler处理的数值特征会带上前缀scaler__,格式为scaler__原始数值特征名
  • 通过remainder='passthrough'保留的特征会直接使用原始名称,不带前缀

方法2:兼容旧版scikit-learn(<0.23)

If you're stuck on an older version, you'll need to manually collect feature names from each transformer and combine them:

# 先拟合你的ColumnTransformer
ct.fit(X)

# 1. 获取独热编码后的特征名称
ohe_transformer = ct.named_transformers_['onehot']
ohe_names = ohe_transformer.get_feature_names_out(cat_features)
# 添加上自定义前缀
ohe_names = [f"onehot__{name}" for name in ohe_names]

# 2. 获取缩放后的数值特征名称
scaler_names = [f"scaler__{name}" for name in num_features]

# 3. 获取remainder保留的特征名称
# 筛选出原始数据中不在cat_features和num_features里的列
remainder_names = [col for col in X.columns if col not in cat_features + num_features]

# 拼接所有特征名称
processed_feature_names = ohe_names + scaler_names + remainder_names

print(processed_feature_names)

关键注意事项:

  • 不管用哪种方法,必须先调用ct.fit(X),因为OneHotEncoder需要先学习数据中的类别才能生成对应的特征名称
  • 如果你的remainder参数不是'passthrough'(比如'drop'),可以跳过拼接remainder特征的步骤

内容的提问来源于stack exchange,提问作者AndreasInfo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.01 02:27:33