You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何识别Pandas DataFrame列中连续NaN并删除超标列?

解决方案

1. 统计每列连续NaN的最大数量

我们可以通过分组统计的方式计算每列中连续NaN的最长序列:

import pandas as pd

nan = float('nan')
data = {'col1': [1, nan, nan, nan, nan, 1, nan, nan], 
        'col2': [1, 1, nan, 1, 0, 0, 1, 0], 
        'col3': [nan, 0, nan, 1, 0, nan, nan, nan], 
        'col4': [1, 0, 0, 1, 0, 1, 1, 1]}
df = pd.DataFrame(data)

# 统计每列连续NaN的最大长度
df_nulls = {}
for col in df.columns:
    # 标记当前列的NaN位置
    is_null = df[col].isna()
    # 生成分组键:非NaN值会重置分组,让连续NaN归为同一组
    group_keys = (~is_null).cumsum()
    # 计算每组的NaN数量,取最大值;无NaN则为0
    max_consec = is_null.groupby(group_keys).sum().max() if is_null.any() else 0
    df_nulls[col] = int(max_consec)

print(df_nulls)
# 输出: {'col1': 4, 'col2': 0, 'col3': 3, 'col4': 0}

如果处理大数据集,也可以用更高效的向量化写法:

max_consec_null = (
    df.isna()
    .cumsum()
    .mask(~df.isna())
    .apply(lambda x: x.groupby(x.notna().cumsum()).count().max())
    .fillna(0)
    .astype(int)
)
df_nulls = max_consec_null.to_dict()

2. 筛选保留符合条件的列

根据统计结果,保留最大连续NaN数不超过2的列:

# 筛选列名
keep_cols = [col for col, cnt in df_nulls.items() if cnt <= 2]
df_filtered = df[keep_cols]

print(df_filtered)

执行后得到的df_filtered即为仅保留col2和col4的目标DataFrame。

内容的提问来源于stack exchange,提问作者Rajesh Ahir

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 09:05:20