如何删除DataFrame中含NaN数量最多的行?
高效删除含NaN数量最多的行/列(循环清除NaN)
核心需求
需要一种高效方法:
- 优先删除含NaN数量最多的行(若有多行并列最多,一并删除)
- 再以同样逻辑删除含NaN最多的列
- 重复上述两步,直到表格中无NaN存在
- 目标:仅通过整行/整列删除操作,保留尽可能多的数据以清除NaN
示例数据
先创建示例DataFrame:
import numpy as np import pandas as pd df = pd.DataFrame([ [1, np.nan, 1, np.nan], [1, 1, 1, 1], [1, np.nan, 1, 1], [np.nan, 1, 1, 1] ], columns=list('ABCD'))
输出结果:
A B C D 0 1.0 NaN 1 NaN 1 1.0 1.0 1 1.0 2 1.0 NaN 1 1.0 3 NaN 1.0 1 1.0
分步解决方案
1. 删除含NaN最多的行
通过向量化操作计算每行的NaN数量,定位并删除所有NaN数量最多的行:
# 计算每行的NaN数量 row_nan_counts = df.isna().sum(axis=1) # 获取最大NaN数量 max_row_nan = row_nan_counts.max() # 筛选出所有NaN数量等于最大值的行索引 rows_to_drop = row_nan_counts[row_nan_counts == max_row_nan].index # 删除目标行 df = df.drop(rows_to_drop)
处理后示例df的结果:
A B C D 1 1.0 1.0 1 1.0 2 1.0 NaN 1 1.0 3 NaN 1.0 1 1.0
2. 删除含NaN最多的列
逻辑与处理行一致,仅调整axis参数:
# 计算每列的NaN数量 col_nan_counts = df.isna().sum(axis=0) # 获取最大NaN数量 max_col_nan = col_nan_counts.max() # 筛选出所有NaN数量等于最大值的列索引 cols_to_drop = col_nan_counts[col_nan_counts == max_col_nan].index # 删除目标列 df = df.drop(cols_to_drop, axis=1)
处理后示例df的结果:
C D 1 1 1.0 2 1 1.0 3 1 1.0
3. 循环清除所有NaN
将上述两步封装为循环,直到表格中无NaN:
import numpy as np import pandas as pd # 初始化示例DataFrame df = pd.DataFrame([ [1, np.nan, 1, np.nan], [1, 1, 1, 1], [1, np.nan, 1, 1], [np.nan, 1, 1, 1] ], columns=list('ABCD')) while df.isna().any().any(): # 处理行:删除含NaN最多的行 row_nan_counts = df.isna().sum(axis=1) max_row_nan = row_nan_counts.max() if max_row_nan > 0: rows_to_drop = row_nan_counts[row_nan_counts == max_row_nan].index df = df.drop(rows_to_drop) # 若删除后已无NaN,提前退出循环 if not df.isna().any().any(): break # 处理列:删除含NaN最多的列 col_nan_counts = df.isna().sum(axis=0) max_col_nan = col_nan_counts.max() if max_col_nan > 0: cols_to_drop = col_nan_counts[col_nan_counts == max_col_nan].index df = df.drop(cols_to_drop, axis=1)
最终结果:
C 1 1 2 1 3 1
方法优势
- 高效性:使用Pandas向量化方法
isna().sum(),避免逐行/列循环,处理大数据集时性能更优 - 灵活性:无需提前指定保留的非NaN数量,动态定位需要删除的行/列
- 彻底性:循环处理直到所有NaN被清除,同时尽可能保留更多数据
内容的提问来源于stack exchange,提问作者dawid
相关产品推荐
相关产品推荐

