You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中数值型数据集缺失值处理:删除或填充的选择及其他技巧

Hey there! Handling missing numerical values is a critical step in data preprocessing, and choosing the right approach depends heavily on your data's characteristics and downstream goals. Let's walk through how to decide between dropping NaNs vs. mean/median imputation, plus some other powerful techniques you can use.

When to Drop NaN Values vs. Use Mean/Median Imputation

Dropping NaNs: Best For Specific Cases

  • Low missing rate (<5%) + Missing Completely at Random (MCAR):If only a tiny fraction of your data is missing, and the missingness has no relation to other features or your target variable, dropping those rows won't hurt your dataset's representativeness. For example, if you have 10k rows and only 300 missing a single numerical feature, deleting those rows is a quick, low-risk move.
    • Quick code snippet:
      df_clean = df.dropna(subset=['your_numerical_col'])
      
  • Irrecoverable non-random missingness:If you can't figure out why values are missing (e.g., missing income data only for high-earners) and keeping those rows would skew your dataset, dropping might be the safest bet—just be sure to document this choice, as it can introduce selection bias.

Mean/Median Imputation: Go-To for Moderate Missingness

  • Mean for normally distributed features:If your feature follows a bell curve, the mean is a solid measure of central tendency. It works well when missing rates are between 5-20% and missingness is random.
  • Median for skewed data or features with outliers:If your data has extreme values (like salary data with a few millionaires), the mean gets pulled toward those outliers. The median is robust to this, so it's a better choice here.
    • Code examples:
      # Mean imputation
      df['col_mean_imputed'] = df['your_numerical_col'].fillna(df['your_numerical_col'].mean())
      # Median imputation
      df['col_median_imputed'] = df['your_numerical_col'].fillna(df['your_numerical_col'].median())
      
  • Warning:If missingness isn't random (e.g., missing temperature readings only on rainy days), simple mean/median imputation will hide that pattern. You'll need more advanced methods here.

Other Missing Value Handling Techniques

Model-Based Imputation

Use other features to predict missing values. This leverages relationships between variables, leading to more accurate fills than simple mean/median. For example, you can use K-Nearest Neighbors (KNN) to fill missing values with the average of similar rows, or train a regression model on complete data to predict missing entries.

  • KNN Imputer example with scikit-learn:
    from sklearn.impute import KNNImputer
    imputer = KNNImputer(n_neighbors=5)  # Use 5 closest rows to fill missing values
    df_imputed = pd.DataFrame(imputer.fit_transform(df), columns=df.columns)
    

Interpolation (Great for Time Series)

If your data is sequential (like daily sales or hourly temperature readings), interpolation uses neighboring values to fill gaps. Linear interpolation is simple and works well for linear trends, while spline interpolation creates smoother curves for non-linear data.

  • Code examples:
    # Linear interpolation
    df['time_series_col'] = df['time_series_col'].interpolate(method='linear')
    # Cubic spline interpolation for smoother fills
    df['time_series_col'] = df['time_series_col'].interpolate(method='spline', order=3)
    

Flag Missing Values + Placeholder Fill

Instead of deleting or imputing, create a new binary feature (e.g., is_missing) to mark when a value was missing, then fill the original column with a unique placeholder (like -999). This lets your model learn if missingness itself is a useful signal—for example, in healthcare data, a missing blood pressure reading might mean the patient wasn't stable enough to take it.

  • Code snippet:
    df['col_is_missing'] = df['your_numerical_col'].isna().astype(int)
    df['your_numerical_col'] = df['your_numerical_col'].fillna(-999)
    

Multiple Imputation

Instead of filling missing values once, generate several versions of your dataset with different imputations (based on statistical distributions), model each one, then combine the results. This accounts for the uncertainty in missing values, leading to more robust conclusions. Tools like fancyimpute's Iterative Imputer (MICE) are popular for this.

  • Example:
    from fancyimpute import IterativeImputer
    imputer = IterativeImputer()  # Uses MICE (Multiple Imputation by Chained Equations)
    df_imputed_list = []
    for _ in range(5):  # Generate 5 imputed datasets
        df_imputed = pd.DataFrame(imputer.fit_transform(df), columns=df.columns)
        df_imputed_list.append(df_imputed)
    # Then model each dataset and average results
    

内容的提问来源于stack exchange,提问作者Kalpit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:51:23