You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scikit Learn决策树中非必填特征的处理方法咨询

Handling Missing Conditional Features in Scikit-Learn Decision Trees

Great question—this is such a common real-world data problem, and you’re already thinking in the right direction by avoiding mindless imputation. Let’s break down a few practical, Scikit-Learn-friendly approaches that let you keep both the "did they resubmit?" signal and the actual resubmission score when it exists:

1. Dual Feature Encoding (Your Initial Idea, Supercharged)

You’re spot-on with creating a HW_RESUBMITTED binary flag (1 if resubmitted, 0 if not). To retain the resubmission score’s value, pair this flag with a filled version of HW_RESUBMISSION where missing values are replaced with a distinct, out-of-range placeholder (since your valid scores are 0-100, something like -1 works perfectly).

Decision trees excel at splitting on these kinds of features: they’ll first learn to split on HW_RESUBMITTED to separate resubmitters from non-resubmitters, then split on the filled score only for the subset of students who did resubmit. No fake interpolated data required!

Here’s how to implement this in pandas:

import pandas as pd

# Create binary flag for resubmission status
df['HW_RESUBMITTED'] = df['HW_RESUBMISSION'].notna().astype(int)
# Fill missing resubmission scores with a unique placeholder
df['HW_RESUBMISSION_FILLED'] = df['HW_RESUBMISSION'].fillna(-1)

You’d then use both HW_RESUBMITTED and HW_RESUBMISSION_FILLED as features in your decision tree.

2. Create a "Resubmission Improvement" Feature

If the difference between the original grade and resubmitted grade is meaningful (e.g., students who improved by 20+ points have different outcomes), you can add this as a derived feature alongside the binary flag. This condenses two pieces of info into one that might have stronger predictive power:

# Calculate score improvement only for resubmitters
df['RESUBMISSION_IMPROVEMENT'] = df['HW_RESUBMISSION'] - df['HW_GRADE']
# Fill missing improvements with an out-of-range value (e.g., -101, since max possible drop is -100)
df['RESUBMISSION_IMPROVEMENT'] = df['RESUBMISSION_IMPROVEMENT'].fillna(-101)

Again, pair this with HW_RESUBMITTED to ensure the model doesn’t confuse "no resubmission" with a huge negative improvement.

3. Lean Into Scikit-Learn’s Native Handling for Placeholder Values

Unlike some models, decision trees don’t care about "invalid" values like -1 as long as they’re consistent. The tree will treat the placeholder as a separate category during splitting—so it’ll naturally learn that students with HW_RESUBMISSION_FILLED = -1 belong to the non-resubmitted group, and split accordingly. You don’t need any extra preprocessing beyond the steps above.

Key Note: Stick With Placeholders, Not Interpolation

You were absolutely right to skip imputation (like mean/median filling)—that would introduce fake data and dilute the true signal. Using a placeholder preserves the missingness as a meaningful category, which decision trees can leverage effectively without distorting your dataset.

内容的提问来源于stack exchange,提问作者pitosalas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:32:05