特征选择后仅返回列索引,如何获取对应DataFrame列名?
Got it, let's fix this issue where you're getting index numbers like [0, 6, 23] instead of actual column names. The root problem lies in your remove_q_qc_dpl_feat function—you're losing the original column names when you convert to numpy arrays and back to DataFrames. Here's how to resolve it:
Why This Happens
When you use VarianceThreshold.transform(), it returns a numpy array (not a DataFrame), which strips away all column metadata. Later, when you transpose and convert back to a DataFrame, it uses default integer indices instead of your original column names. By the time you run SelectPercentile, your X_train no longer has the original column labels, so get_support() returns positions instead of names.
Fixed remove_q_qc_dpl_feat Function
Modify the function to preserve original column names throughout the process:
import pandas as pd from sklearn.feature_selection import VarianceThreshold def remove_q_qc_dpl_feat(X_train, X_test, y_train, y_test): # Remove constant and quasi-constant features (keep column names) constant_filter = VarianceThreshold(threshold=0.01) constant_filter.fit(X_train) # Get boolean mask of features to keep, then filter original DataFrame keep_constant = constant_filter.get_support() X_train_filter = X_train.loc[:, keep_constant] X_test_filter = X_test.loc[:, keep_constant] # Remove duplicate features (keep original column names) # Transpose to check duplicate rows (which correspond to original columns) X_train_T = X_train_filter.T # Get mask of non-duplicated rows (original columns) not_duplicated = ~X_train_T.duplicated() # Filter and transpose back to original shape X_train_unique = X_train_T[not_duplicated].T X_test_unique = X_test_filter.loc[:, X_train_unique.columns] return X_train_unique, X_test_unique
How This Works
- Instead of using
transform()(which outputs a numpy array), we useget_support()to get a boolean mask, then filter the original DataFrame directly. This keeps all column names intact. - For duplicate features, we transpose the filtered DataFrame (so rows represent original columns), check for duplicates, then filter and transpose back—keeping the original column names the entire time.
Updated Calling Code
Now when you run your feature selection code, X_train will still have its original column names, so features will return the actual names instead of indices:
X_train, X_test = remove_q_qc_dpl_feat(X_train, X_test, y_train, y_test) sel = SelectPercentile(mutual_info_classif, percentile=10).fit(X_train, y_train) # No need to call fit twice—you already did it above! features = X_train.columns[sel.get_support()] print(features) # This will now show your original column names, not indices
Note: I removed the duplicate sel.fit(X_train, y_train) call—you only need to fit once!
内容的提问来源于stack exchange,提问作者dJudge

