You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除DataFrame中含NaN数量最多的行?

高效删除含NaN数量最多的行/列(循环清除NaN)

核心需求

需要一种高效方法:

  • 优先删除含NaN数量最多的行(若有多行并列最多,一并删除)
  • 再以同样逻辑删除含NaN最多的列
  • 重复上述两步,直到表格中无NaN存在
  • 目标:仅通过整行/整列删除操作,保留尽可能多的数据以清除NaN

示例数据

先创建示例DataFrame:

import numpy as np
import pandas as pd

df = pd.DataFrame([
    [1, np.nan, 1, np.nan],
    [1, 1, 1, 1],
    [1, np.nan, 1, 1],
    [np.nan, 1, 1, 1]
], columns=list('ABCD'))

输出结果:

A    B  C    D
0  1.0  NaN  1  NaN
1  1.0  1.0  1  1.0
2  1.0  NaN  1  1.0
3  NaN  1.0  1  1.0

分步解决方案

1. 删除含NaN最多的行

通过向量化操作计算每行的NaN数量,定位并删除所有NaN数量最多的行:

# 计算每行的NaN数量
row_nan_counts = df.isna().sum(axis=1)
# 获取最大NaN数量
max_row_nan = row_nan_counts.max()
# 筛选出所有NaN数量等于最大值的行索引
rows_to_drop = row_nan_counts[row_nan_counts == max_row_nan].index
# 删除目标行
df = df.drop(rows_to_drop)

处理后示例df的结果:

A    B  C    D
1  1.0  1.0  1  1.0
2  1.0  NaN  1  1.0
3  NaN  1.0  1  1.0

2. 删除含NaN最多的列

逻辑与处理行一致,仅调整axis参数:

# 计算每列的NaN数量
col_nan_counts = df.isna().sum(axis=0)
# 获取最大NaN数量
max_col_nan = col_nan_counts.max()
# 筛选出所有NaN数量等于最大值的列索引
cols_to_drop = col_nan_counts[col_nan_counts == max_col_nan].index
# 删除目标列
df = df.drop(cols_to_drop, axis=1)

处理后示例df的结果:

C    D
1  1  1.0
2  1  1.0
3  1  1.0

3. 循环清除所有NaN

将上述两步封装为循环,直到表格中无NaN:

import numpy as np
import pandas as pd

# 初始化示例DataFrame
df = pd.DataFrame([
    [1, np.nan, 1, np.nan],
    [1, 1, 1, 1],
    [1, np.nan, 1, 1],
    [np.nan, 1, 1, 1]
], columns=list('ABCD'))

while df.isna().any().any():
    # 处理行:删除含NaN最多的行
    row_nan_counts = df.isna().sum(axis=1)
    max_row_nan = row_nan_counts.max()
    if max_row_nan > 0:
        rows_to_drop = row_nan_counts[row_nan_counts == max_row_nan].index
        df = df.drop(rows_to_drop)
        # 若删除后已无NaN,提前退出循环
        if not df.isna().any().any():
            break
    
    # 处理列:删除含NaN最多的列
    col_nan_counts = df.isna().sum(axis=0)
    max_col_nan = col_nan_counts.max()
    if max_col_nan > 0:
        cols_to_drop = col_nan_counts[col_nan_counts == max_col_nan].index
        df = df.drop(cols_to_drop, axis=1)

最终结果:

C
1  1
2  1
3  1

方法优势

  • 高效性:使用Pandas向量化方法isna().sum(),避免逐行/列循环,处理大数据集时性能更优
  • 灵活性:无需提前指定保留的非NaN数量,动态定位需要删除的行/列
  • 彻底性:循环处理直到所有NaN被清除,同时尽可能保留更多数据

内容的提问来源于stack exchange,提问作者dawid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 03:42:06