Scikit Learn决策树中非必填特征的处理方法咨询
Great question—this is such a common real-world data problem, and you’re already thinking in the right direction by avoiding mindless imputation. Let’s break down a few practical, Scikit-Learn-friendly approaches that let you keep both the "did they resubmit?" signal and the actual resubmission score when it exists:
1. Dual Feature Encoding (Your Initial Idea, Supercharged)
You’re spot-on with creating a HW_RESUBMITTED binary flag (1 if resubmitted, 0 if not). To retain the resubmission score’s value, pair this flag with a filled version of HW_RESUBMISSION where missing values are replaced with a distinct, out-of-range placeholder (since your valid scores are 0-100, something like -1 works perfectly).
Decision trees excel at splitting on these kinds of features: they’ll first learn to split on HW_RESUBMITTED to separate resubmitters from non-resubmitters, then split on the filled score only for the subset of students who did resubmit. No fake interpolated data required!
Here’s how to implement this in pandas:
import pandas as pd # Create binary flag for resubmission status df['HW_RESUBMITTED'] = df['HW_RESUBMISSION'].notna().astype(int) # Fill missing resubmission scores with a unique placeholder df['HW_RESUBMISSION_FILLED'] = df['HW_RESUBMISSION'].fillna(-1)
You’d then use both HW_RESUBMITTED and HW_RESUBMISSION_FILLED as features in your decision tree.
2. Create a "Resubmission Improvement" Feature
If the difference between the original grade and resubmitted grade is meaningful (e.g., students who improved by 20+ points have different outcomes), you can add this as a derived feature alongside the binary flag. This condenses two pieces of info into one that might have stronger predictive power:
# Calculate score improvement only for resubmitters df['RESUBMISSION_IMPROVEMENT'] = df['HW_RESUBMISSION'] - df['HW_GRADE'] # Fill missing improvements with an out-of-range value (e.g., -101, since max possible drop is -100) df['RESUBMISSION_IMPROVEMENT'] = df['RESUBMISSION_IMPROVEMENT'].fillna(-101)
Again, pair this with HW_RESUBMITTED to ensure the model doesn’t confuse "no resubmission" with a huge negative improvement.
3. Lean Into Scikit-Learn’s Native Handling for Placeholder Values
Unlike some models, decision trees don’t care about "invalid" values like -1 as long as they’re consistent. The tree will treat the placeholder as a separate category during splitting—so it’ll naturally learn that students with HW_RESUBMISSION_FILLED = -1 belong to the non-resubmitted group, and split accordingly. You don’t need any extra preprocessing beyond the steps above.
Key Note: Stick With Placeholders, Not Interpolation
You were absolutely right to skip imputation (like mean/median filling)—that would introduce fake data and dilute the true signal. Using a placeholder preserves the missingness as a meaningful category, which decision trees can leverage effectively without distorting your dataset.
内容的提问来源于stack exchange,提问作者pitosalas

