如何对离散变量进行缺失值插补?含非众数插补(计数变量)方案
Awesome question! Handling missing values for discrete variables is a super common pain point in data preprocessing, and avoiding mode imputation for count-based discrete variables adds a nice layer of nuance. Let’s break this down clearly.
General Imputation Methods for Discrete Variables
These work for categorical discrete variables (like gender, occupation) and can sometimes apply to count variables too:
- Category Labeling: Treat missing values as a brand-new category (e.g., label missing gender as "Unknown"). This preserves the information that the value was missing (which might be meaningful—like users refusing to share demographic data) and is dead simple to implement.
- Model-Based Predictive Imputation: Train a classification model (e.g., random forest classifier, logistic regression) using other non-missing features to predict the missing discrete values. For example, if you’re missing occupation data, use age, income, and location as features to train a model that guesses the missing occupation.
- Hot-Deck Imputation: Find non-missing samples that are statistically similar to the missing one (using features like age, region) and borrow their discrete value. Think of it as "filling in with a peer’s data."
- Cold-Deck Imputation: Similar to hot-deck, but instead of using your current dataset, you pull values from an external dataset with a similar category distribution. Great if your current dataset is too small to find good matches.
Count-Based Discrete Variables: Imputation Without Mode
Count variables (like number of purchases, support tickets opened) are discrete but have numerical properties—here are methods that avoid leaning on the mode (which can skew results if the mode is an overrepresented value like 0):
- Poisson/Negative Binomial Regression Imputation: Count data often follows Poisson or negative binomial distributions. Train a regression model using related features (e.g., user tenure, membership level) to predict the missing count values. This method respects the natural distribution of your count data, which mode imputation doesn’t.
- Stratified Mean Imputation: Split your data into subgroups based on relevant features (e.g., split customers by membership tier), calculate the mean count for each subgroup, then fill missing values with the rounded mean of their subgroup. This is way more targeted than using a global mean, and avoids the mode’s bias.
- Multiple Imputation (MICE): Use Multiple Imputation by Chained Equations (MICE) with a count-specific model (like Poisson MICE). This generates several plausible imputed datasets, accounts for uncertainty in the imputation, and is ideal for rigorous statistical analysis. You’ll analyze all datasets and combine results to reduce bias.
- Zero-Inflated Model Imputation: If your count data has tons of 0s (e.g., number of monthly app logins where most users don’t log in), use a zero-inflated Poisson or negative binomial model. First predict whether a missing value should be 0, then predict the non-zero count if applicable—this handles the unique distribution of zero-heavy count data perfectly.
- Distribution-Based Random Sampling: First estimate the distribution parameters of your non-missing count data (e.g., calculate λ for a Poisson distribution). Then randomly sample values from this distribution to fill missing entries. This preserves the original data’s distribution instead of forcing all missing values into the most common count.
A quick reminder: Always pick a method that fits your data’s context and your end goal. If you’re building a predictive model, make sure to do your imputation only on the training data (not the test set) to avoid data leakage!
内容的提问来源于stack exchange,提问作者Gale

