如何让逻辑与运算符(&)按另一侧值处理NaN?Pandas/Numpy高效方案
自定义布尔与运算(含NaN)的Pandas优化实现
需求定义
需实现自定义与运算规则,处理含NaN的布尔值:
True & NaN = TrueFalse & NaN = FalseNaN & NaN = NaN
现有实现方式
初始实现代码:
(a.fillna(True) & b.fillna(True)).where(~(a.isna() & b.isna()), None)
示例验证
示例代码
from itertools import product import pandas as pd from IPython.display import display a = pd.DataFrame((product([True, False, None], [True, False, None]))) display(a) display((a[0].fillna(True) & a[1].fillna(True)).where(~(a[0].isna() & a[1].isna()), None))
输出结果
0 1 0 True True 1 True False 2 True None 3 False True 4 False False 5 False None 6 None True 7 None False 8 None None 0 True 1 False 2 True 3 False 4 False 5 False 6 True 7 False 8 None dtype: object
场景化最优实现
针对两种数据分布场景,分别推荐最优方案:
场景A:仅少量行包含全NaN
推荐方案:利用all()结合mask()
b.all(1).mask(b.isna().all(1))
性能表现:100次循环耗时约2.4s(基于10万行样本,仅117行全NaN)
场景B:多数行包含全NaN
推荐方案:利用stack()+groupby()+reindex()
c.stack().groupby(level=0).all().reindex(c.index)
性能表现:100次循环耗时约0.9s(基于10万行样本,约90879行全NaN)
性能测试细节
测试代码
# 构造测试样本 b = a.sample(int(1e5), weights=[1,1,1,1,1,1,1,1,0.01], ignore_index=True, replace=True) c = a.sample(int(1e5), weights=[1,1,1,1,1,1,1,1,80], ignore_index=True, replace=True) print(b.isna().all(axis="columns").sum()) # 输出:117 full NaN row print(c.isna().all(axis="columns").sum()) # 输出:90879 full NaN rows # 性能计时 import timeit print(timeit.timeit(lambda: b.all(1).mask(b.isna().all(1)), number=100)) # 输出:2.4s print(timeit.timeit(lambda: c.all(1).mask(c.isna().all(1)), number=100)) # 输出:1.6s print(timeit.timeit(lambda: b.stack().groupby(level=0).all().reindex(b.index), number=100)) # 输出:3.3s print(timeit.timeit(lambda: c.stack().groupby(level=0).all().reindex(c.index), number=100)) # 输出:0.9s
原理说明
all()+mask()方案:先按行执行与运算(Pandas默认all()忽略NaN),再通过mask()将全NaN行重置为None。适合全NaN行占比低的场景,无需大量数据重组,计算成本低。stack()方案:先剔除所有NaN数据堆叠,按原行分组计算all(),最后重新索引补全全NaN行为None。当多数行是全NaN时,堆叠后的数据量大幅减少,计算效率显著提升。
内容的提问来源于stack exchange,提问作者Wang
相关产品推荐
相关产品推荐

