You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

特征选择后仅返回列索引,如何获取对应DataFrame列名?

How to Get Original Column Names Instead of Indices After Feature Selection

Got it, let's fix this issue where you're getting index numbers like [0, 6, 23] instead of actual column names. The root problem lies in your remove_q_qc_dpl_feat function—you're losing the original column names when you convert to numpy arrays and back to DataFrames. Here's how to resolve it:

Why This Happens

When you use VarianceThreshold.transform(), it returns a numpy array (not a DataFrame), which strips away all column metadata. Later, when you transpose and convert back to a DataFrame, it uses default integer indices instead of your original column names. By the time you run SelectPercentile, your X_train no longer has the original column labels, so get_support() returns positions instead of names.

Fixed remove_q_qc_dpl_feat Function

Modify the function to preserve original column names throughout the process:

import pandas as pd
from sklearn.feature_selection import VarianceThreshold

def remove_q_qc_dpl_feat(X_train, X_test, y_train, y_test):
    # Remove constant and quasi-constant features (keep column names)
    constant_filter = VarianceThreshold(threshold=0.01)
    constant_filter.fit(X_train)
    # Get boolean mask of features to keep, then filter original DataFrame
    keep_constant = constant_filter.get_support()
    X_train_filter = X_train.loc[:, keep_constant]
    X_test_filter = X_test.loc[:, keep_constant]
    
    # Remove duplicate features (keep original column names)
    # Transpose to check duplicate rows (which correspond to original columns)
    X_train_T = X_train_filter.T
    # Get mask of non-duplicated rows (original columns)
    not_duplicated = ~X_train_T.duplicated()
    # Filter and transpose back to original shape
    X_train_unique = X_train_T[not_duplicated].T
    X_test_unique = X_test_filter.loc[:, X_train_unique.columns]
    
    return X_train_unique, X_test_unique

How This Works

  • Instead of using transform() (which outputs a numpy array), we use get_support() to get a boolean mask, then filter the original DataFrame directly. This keeps all column names intact.
  • For duplicate features, we transpose the filtered DataFrame (so rows represent original columns), check for duplicates, then filter and transpose back—keeping the original column names the entire time.

Updated Calling Code

Now when you run your feature selection code, X_train will still have its original column names, so features will return the actual names instead of indices:

X_train, X_test = remove_q_qc_dpl_feat(X_train, X_test, y_train, y_test)
sel = SelectPercentile(mutual_info_classif, percentile=10).fit(X_train, y_train)
# No need to call fit twice—you already did it above!
features = X_train.columns[sel.get_support()]
print(features)  # This will now show your original column names, not indices

Note: I removed the duplicate sel.fit(X_train, y_train) call—you only need to fit once!

内容的提问来源于stack exchange,提问作者dJudge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 20:37:44