Python Pandas 如何用函数式编程合并多个布尔Series
问题描述
你持有一个包含gender、marital、education等多个列的pandas DataFrame,样例数据结构如下:
| gender | marital | education |
|---|---|---|
| male | single | tertiary |
你已经生成了一组存储在列表中的布尔Series筛选条件,条件生成代码如下:
bool_gender = df["gender"] == "male" bool_marital = df["marital"] == "married" bool_education = df["education"] == "secondary" cond_list = [bool_gender, bool_marital, bool_education]
需求是在Python 3环境下使用函数式编程写法,将列表内所有布尔Series通过&运算符做逻辑与合并,最终得到的单一布尔Series需要和以下表达式的计算结果完全一致:
desired_output = bool_gender & bool_marital & bool_education
预判实现可能会用到类似reduce("&", map(function, iter))的形式。
实现方案
直接使用functools.reduce做累积运算即可,不需要额外嵌套map,注意reduce不能直接传入字符串形式的运算符,需要调用运算符对应的可调用对象:
- 推荐写法:借助标准库
operator的and_方法,性能比lambda更好,逻辑和原生&完全等价
from functools import reduce import operator combined_cond = reduce(operator.and_, cond_list)
注:pandas布尔Series重载了
__and__魔术方法,operator.and_本质就是调用两个Series的&运算逻辑,逐元素返回逻辑与结果,和手动链式拼接条件的输出没有差异。
- 轻量写法:如果不想导入
operator模块,直接传入lambda表达式即可,效果和上面的写法完全一致
from functools import reduce combined_cond = reduce(lambda x, y: x & y, cond_list)
两种写法都属于标准的函数式编程实现,无需编写显式循环,也不需要手动逐行拼接条件,不管列表里有多少个布尔筛选条件都可以正常合并。
内容的提问来源于stack exchange,提问作者matt.aurelio
相关产品推荐
相关产品推荐

