Python中数值型数据集缺失值处理:删除或填充的选择及其他技巧
Hey there! Handling missing numerical values is a critical step in data preprocessing, and choosing the right approach depends heavily on your data's characteristics and downstream goals. Let's walk through how to decide between dropping NaNs vs. mean/median imputation, plus some other powerful techniques you can use.
When to Drop NaN Values vs. Use Mean/Median Imputation
Dropping NaNs: Best For Specific Cases
- Low missing rate (<5%) + Missing Completely at Random (MCAR):If only a tiny fraction of your data is missing, and the missingness has no relation to other features or your target variable, dropping those rows won't hurt your dataset's representativeness. For example, if you have 10k rows and only 300 missing a single numerical feature, deleting those rows is a quick, low-risk move.
- Quick code snippet:
df_clean = df.dropna(subset=['your_numerical_col'])
- Quick code snippet:
- Irrecoverable non-random missingness:If you can't figure out why values are missing (e.g., missing income data only for high-earners) and keeping those rows would skew your dataset, dropping might be the safest bet—just be sure to document this choice, as it can introduce selection bias.
Mean/Median Imputation: Go-To for Moderate Missingness
- Mean for normally distributed features:If your feature follows a bell curve, the mean is a solid measure of central tendency. It works well when missing rates are between 5-20% and missingness is random.
- Median for skewed data or features with outliers:If your data has extreme values (like salary data with a few millionaires), the mean gets pulled toward those outliers. The median is robust to this, so it's a better choice here.
- Code examples:
# Mean imputation df['col_mean_imputed'] = df['your_numerical_col'].fillna(df['your_numerical_col'].mean()) # Median imputation df['col_median_imputed'] = df['your_numerical_col'].fillna(df['your_numerical_col'].median())
- Code examples:
- Warning:If missingness isn't random (e.g., missing temperature readings only on rainy days), simple mean/median imputation will hide that pattern. You'll need more advanced methods here.
Other Missing Value Handling Techniques
Model-Based Imputation
Use other features to predict missing values. This leverages relationships between variables, leading to more accurate fills than simple mean/median. For example, you can use K-Nearest Neighbors (KNN) to fill missing values with the average of similar rows, or train a regression model on complete data to predict missing entries.
- KNN Imputer example with scikit-learn:
from sklearn.impute import KNNImputer imputer = KNNImputer(n_neighbors=5) # Use 5 closest rows to fill missing values df_imputed = pd.DataFrame(imputer.fit_transform(df), columns=df.columns)
Interpolation (Great for Time Series)
If your data is sequential (like daily sales or hourly temperature readings), interpolation uses neighboring values to fill gaps. Linear interpolation is simple and works well for linear trends, while spline interpolation creates smoother curves for non-linear data.
- Code examples:
# Linear interpolation df['time_series_col'] = df['time_series_col'].interpolate(method='linear') # Cubic spline interpolation for smoother fills df['time_series_col'] = df['time_series_col'].interpolate(method='spline', order=3)
Flag Missing Values + Placeholder Fill
Instead of deleting or imputing, create a new binary feature (e.g., is_missing) to mark when a value was missing, then fill the original column with a unique placeholder (like -999). This lets your model learn if missingness itself is a useful signal—for example, in healthcare data, a missing blood pressure reading might mean the patient wasn't stable enough to take it.
- Code snippet:
df['col_is_missing'] = df['your_numerical_col'].isna().astype(int) df['your_numerical_col'] = df['your_numerical_col'].fillna(-999)
Multiple Imputation
Instead of filling missing values once, generate several versions of your dataset with different imputations (based on statistical distributions), model each one, then combine the results. This accounts for the uncertainty in missing values, leading to more robust conclusions. Tools like fancyimpute's Iterative Imputer (MICE) are popular for this.
- Example:
from fancyimpute import IterativeImputer imputer = IterativeImputer() # Uses MICE (Multiple Imputation by Chained Equations) df_imputed_list = [] for _ in range(5): # Generate 5 imputed datasets df_imputed = pd.DataFrame(imputer.fit_transform(df), columns=df.columns) df_imputed_list.append(df_imputed) # Then model each dataset and average results
内容的提问来源于stack exchange,提问作者Kalpit

