You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据框缺失值处理:用for循环删除高缺失特征并均值填充

处理数据框缺失值的For循环实现

用Python的pandas库可以轻松实现需求,以下是完整代码和步骤说明:

import pandas as pd

# 假设你的目标数据框为df
# 先计算每列的缺失值占比
null_proportions = df.isnull().sum() / len(df)

# 遍历每一列做缺失值处理
for column in df.columns:
    prop = null_proportions[column]
    if prop >= 0.2:
        # 删除缺失值占比≥20%的特征
        df.drop(column, axis=1, inplace=True)
    elif prop > 0:
        # 用特征均值填充缺失值占比<20%的特征
        df[column].fillna(df[column].mean(), inplace=True)

代码说明

  • 计算缺失占比:df.isnull().sum()统计每列的缺失值总数,除以数据框总行数len(df)得到每列的缺失值占比
  • 循环判断处理:
    • 当某列缺失占比≥20%时,调用drop方法删除该列(axis=1指定按列操作,inplace=True直接修改原数据框)
    • 当某列存在缺失但占比<20%时,用该列的均值(df[column].mean())填充缺失值
    • 无缺失值的列会自动跳过处理

内容的提问来源于stack exchange,提问作者Vidhi Bafna

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 15:35:28