报错:'DecisionTreeClassifier'对象无'transform'属性,求解决方法
Got it, let's break down why you're hitting this error and how to fix it quickly.
Why This Error Occurs
The core issue here is that DecisionTreeClassifier is a classification model, not a feature selection transformer. It doesn't come with a built-in transform() method. Your code tries to use it for entropy-based feature selection by chaining .fit(X, y).transform(X), but that's not what the DT classifier is designed to do.
When you fit a DecisionTreeClassifier, it does calculate feature importance scores based on entropy, but it can't directly "transform" your dataset to keep only the important features. That's a job for dedicated feature selection tools in scikit-learn.
Fixes to Resolve the Error
We have two simple, scikit-learn-native ways to fix this, both leveraging the DT's feature importance scores for selection:
Option 1: Use SelectFromModel (Cleanest Approach)
SelectFromModel is a transformer that wraps any estimator with a feature_importances_ attribute (like DecisionTreeClassifier) to automatically select features based on importance thresholds. Here's how to update your feature selection block:
from sklearn.feature_selection import SelectFromModel # Replace your existing feature selection code with this if doFeatureSelection: print 'Performing Feature Selection:' print 'Shape of dataset before feature selection: ' + str(X.shape) # Wrap the DT classifier in SelectFromModel selector = SelectFromModel(DecisionTreeClassifier(criterion='entropy')) X = selector.fit_transform(X, y) print 'Shape of dataset after feature selection: ' + str(X.shape) + '\n'
By default, this keeps features whose importance is above the mean of all importances. You can tweak the threshold with the threshold parameter (e.g., threshold='median' or a numeric value like 0.05) if you want more control over how many features are retained.
Option 2: Manually Filter Features (Full Control)
If you want to set a custom threshold or have more visibility into which features are kept, you can extract the importance scores from the fitted DT and filter the features yourself:
if doFeatureSelection: print 'Performing Feature Selection:' print 'Shape of dataset before feature selection: ' + str(X.shape) clf = DecisionTreeClassifier(criterion='entropy') clf.fit(X, y) # Get feature importances and set a threshold (e.g., mean importance) importances = clf.feature_importances_ threshold = importances.mean() # Create a mask to keep only features above the threshold mask = importances > threshold X = X[:, mask] print 'Shape of dataset after feature selection: ' + str(X.shape) + '\n'
This lets you adjust the threshold to be higher (fewer features) or lower (more features) based on your needs.
Quick Side Note
If you're using Python 3, remember that print is a function, so you should use parentheses (e.g., print('Encoding feature...')) instead of the Python 2 syntax in your code. No big deal if you're still on Python 2, but it's a heads-up for future upgrades.
内容的提问来源于stack exchange,提问作者Aparna

