如何在Pandas中移除指定列含NaN/None或空列表的行?
解决Pandas移除指定列全为NaN/None/空列表的行的问题
原始数据
首先定义目标DataFrame:
import pandas as pd df = pd.DataFrame({'animal': ['Falcon', 'Falcon', 'Parrot', 'Parrot'], 'speed': [380., 370., 24., None], 'height': [380., 370., [], None], 'weight': [380., 370., None, None]})
对应的表格输出:
| animal | speed | height | weight | |
|---|---|---|---|---|
| 0 | Falcon | 380.0 | 380.0 | 380.0 |
| 1 | Falcon | 370.0 | 370.0 | 370.0 |
| 2 | Parrot | 24.0 | [] | NaN |
| 3 | Parrot | NaN | None | NaN |
需求
移除height和weight列中**所有值均为NaN/None或空列表[]**的行。
原有方法的局限
使用dropna只能识别NaN/None,无法处理空列表,执行以下代码后,第2行(含空列表)会被保留:
cols = ['height','weight'] df2 = df.dropna(axis=0, how='all', subset=cols)
得到的结果:
| animal | speed | height | weight | |
|---|---|---|---|---|
| 0 | Falcon | 380.0 | 380.0 | 380.0 |
| 1 | Falcon | 370.0 | 370.0 | 370.0 |
| 2 | Parrot | 24.0 | [] | NaN |
解决办法
方法一:自定义空值判断(通用型)
先定义函数识别所有空值类型(NaN/None/空列表),再标记出指定列全为空的行,最后过滤:
cols = ['height','weight'] # 判断值是否为需要处理的空值 def is_empty(val): return pd.isna(val) or (isinstance(val, list) and len(val) == 0) # 生成掩码:指定列全为空的行标记为True mask = df[cols].applymap(is_empty).all(axis=1) # 保留非全空的行 df_result = df[~mask]
方法二:转换空列表为缺失值(高效型)
将空列表统一转为Pandas缺失值pd.NA,再复用dropna逻辑,适合大数据量场景:
cols = ['height','weight'] # 遍历指定列,把空列表替换为pd.NA for col in cols: df[col] = df[col].apply(lambda x: pd.NA if isinstance(x, list) and len(x) == 0 else x) # 移除指定列全为缺失值的行 df_result = df.dropna(axis=0, how='all', subset=cols)
两种方法最终得到的结果一致:
| animal | speed | height | weight | |
|---|---|---|---|---|
| 0 | Falcon | 380.0 | 380.0 | 380.0 |
| 1 | Falcon | 370.0 | 370.0 | 370.0 |
内容的提问来源于stack exchange,提问作者user173729
相关产品推荐
相关产品推荐

