You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于含NaN通配符的Series高效过滤Pandas DataFrame?

Pandas DataFrame按带NaN通配符的Series过滤的优化方案

需求回顾

需要用一个Series作为过滤条件筛选DataFrame,其中Series中的NaN作为通配符,可匹配任意值;原实现用iterrows()逐行遍历,效率低且代码不够简洁。

优化思路

核心是抛弃逐行遍历,利用Pandas的列级向量化操作实现过滤——这是Pandas处理大数据集的最优方式,底层基于C语言实现,效率远高于Python级别的循环。

具体实现步骤

  1. 提取过滤条件Series中非NaN的条目,因为NaN作为通配符无需参与匹配判断
  2. 对每个非NaN的列,生成DataFrame对应列与过滤值相等的布尔数组
  3. 将所有布尔数组做逻辑与(&)运算,得到最终的过滤掩码
  4. 用掩码直接筛选DataFrame的行

优化后代码

import pandas as pd

if __name__ == '__main__':
    data = {'Name': ['Ankit', 'Amit', 'Aishwarya', 'Priyanka'],
            'Age': [21, 19, 19, 19],
            'Stream': ['Math', 'Commerce', 'Arts', 'Biology'],
            'Percentage': [88, 88, 88, 70]}

    df = pd.DataFrame(data, columns=['Name', 'Age', 'Stream', 'Percentage'])

    # 用pd.NA替代math.nan,更符合Pandas空值规范
    f = pd.Series([pd.NA, 19, pd.NA, 88], index=['Name', 'Age', 'Stream', 'Percentage'])

    # 初始化掩码为全True
    filter_mask = pd.Series([True] * len(df), index=df.index)

    # 遍历非空过滤条件,逐步缩小匹配范围
    for col, val in f.dropna().items():
        filter_mask &= df[col] == val

    filtered_df = df[filter_mask]
    print(filtered_df)

效率优势

  • 原代码的iterrows()是Python级别的逐行循环,当DataFrame行数超过1万时,速度会急剧下降
  • 优化后的代码用列级向量化运算,所有匹配逻辑在Pandas底层执行,效率提升可达10~100倍(取决于数据集大小)

更简洁的替代写法

如果追求代码极致简洁,可利用isin()和all(axis=1)实现,效率与上述方案接近:

# 简洁版实现
filtered_df = df[df.isin(f.dropna().to_dict()).all(axis=1)]

内容的提问来源于stack exchange,提问作者Emil Jansson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.15 16:01:07