如何从数组获取所需列索引并将对应列添加至DataFrame?
Great question! You’re on the right track using get_support(indices=True) to grab selected column indices, but we can make this process cleaner (and more efficient, since looping isn’t ideal for pandas operations) with a straightforward approach.
Optimal Approach (No Loop Needed)
Pandas’ iloc indexer is built for selecting columns by their integer positions. Since RequiredColumns gives you the exact indices of the columns you want, you can slice your original DataFrame in one line:
import pandas as pd # Sample DataFrame matching your structure X = pd.DataFrame({'A': [1,1,1,1], 'B': [2,2,2,2], 'C': [3,3,3,3], 'D': [4,4,4,4]}) RequiredColumns = [0, 1] # Result from skb.get_support(indices=True) # Create the filtered DataFrame in a single step FeatureSelected_DataFrame = X.iloc[:, RequiredColumns]
This will immediately give you a DataFrame with columns A and B (since indices 0 and 1 map to those columns).
Fixing Your Loop-Based Approach (If You Prefer)
If you want to complete your existing loop code (though it’s less efficient for large datasets), you can use pd.concat to append each selected column to your new DataFrame:
FeatureSelected_DataFrame = pd.DataFrame() for i in RequiredColumns: # Select the column at index i (keep it as a DataFrame with [i]) ReqColumn = X.iloc[:, [i]] # Append the column to the new DataFrame (axis=1 adds it as a column) FeatureSelected_DataFrame = pd.concat([FeatureSelected_DataFrame, ReqColumn], axis=1)
Alternatively, you can assign columns directly using their names:
FeatureSelected_DataFrame = pd.DataFrame() for i in RequiredColumns: col_name = X.columns[i] FeatureSelected_DataFrame[col_name] = X[col_name]
Key Notes
iloc[:, RequiredColumns]is the most efficient method here because it uses pandas’ vectorized operations instead of slow row/column-wise iteration.- Ensure
Xis a pandas DataFrame (not a numpy array) for these methods to work. If it’s an array, convert it first withpd.DataFrame(X).
内容的提问来源于stack exchange,提问作者Student of the Digital World

