缺失值插补技术问询:训练集与预测集的取值选择及后续处理
Missing Value Imputation: Key Rules for Training and Prediction
Let’s tackle these two questions one by one—they’re tied together by a core principle that’s easy to mix up, so let’s break it down plainly: never use information from your prediction dataset when training your model, because that’s "data leakage" that will break your model’s ability to perform on real, unseen data.
Question 1: Which mean/median to use for imputation?
Always use the training set’s mean or median when filling in missing values. Here’s the logic:
- When you’re building your model, the prediction dataset (the data you’ll eventually use to make real-world predictions) is completely "unseen" during training. If you calculate stats like mean from the combined training + prediction data, or just the prediction data alone, you’re giving your model a peek at information it wouldn’t have access to in a live scenario.
- This makes your model look great during testing, but it’ll crash and burn when deployed on truly new data.
Question 2: How to impute missing values in the prediction set post-training?
Stick strictly with the exact value you used to impute the training set—in your example, that’s the training set’s paw_number mean of 6.
- Your model learned patterns based on the training data’s feature distribution, where missing paw_number values were replaced with 6. If you swap in a different value (like the prediction set’s own mean) for new data, you’re altering the feature’s distribution in a way the model never learned to handle.
- For instance, if your prediction set’s paw_number mean is 5, replacing missing values with 5 would create a shift the model hasn’t seen, leading to unreliable, inconsistent predictions. Consistency here is non-negotiable to keep your model’s inputs aligned with what it was trained on.
内容的提问来源于stack exchange,提问作者compguy24
相关产品推荐
相关产品推荐

