如何列表存储、拼接并执行Polars过滤表达式?
问题:如何批量拼接Polars过滤器并在
.filter()中生效? 我希望将多个不同的过滤器存储在一个对象(列表、字典等)中,之后可以选择所需过滤器并在Polars的.filter()方法中执行。示例如下:
# Sample DataFrame df = pl.DataFrame( {"col_a": [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], "col_b": [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]} ) # Set a couple of filters filter_1 = pl.col("col_a") > 5 filter_2 = pl.col("col_b") > 8 # Apply filters: this works fine! df_filtered = df.filter(filter_1 & filter_2) # Concatenate filters filters = [filter_1, filter_2] # This won't work: df.filter((" & ").join(filters)) df.filter((" | ").join(filters))
请问如何实现类似(" & ").join(filters)的正确拼接方式,使其能在.filter()方法中生效?
解决方案
1. 逻辑与(AND)拼接过滤器
使用functools.reduce结合Polars的&运算符,可将列表中的过滤器依次组合为逻辑与关系:
from functools import reduce import polars as pl df = pl.DataFrame( {"col_a": [1, 2, 3, 4, 5, 6, 7, 8, 9, 10], "col_b": [1, 2, 3, 4, 5, 6, 7, 8, 9, 10]} ) filter_1 = pl.col("col_a") > 5 filter_2 = pl.col("col_b") > 8 filters = [filter_1, filter_2] # 组合逻辑与过滤器 combined_filter = reduce(lambda a, b: a & b, filters) df_filtered = df.filter(combined_filter) print(df_filtered)
输出结果:
shape: (2, 2) ┌───────┬───────┐ │ col_a ┆ col_b │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞═══════╪═══════╡ │ 9 ┆ 9 │ │ 10 ┆ 10 │ └───────┴───────┘
2. 逻辑或(OR)拼接过滤器
同样用reduce,替换运算符为|即可实现逻辑或的组合:
# 组合逻辑或过滤器 combined_filter = reduce(lambda a, b: a | b, filters) df_filtered = df.filter(combined_filter) print(df_filtered)
输出结果:
shape: (7, 2) ┌───────┬───────┐ │ col_a ┆ col_b │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞═══════╪═══════╡ │ 6 ┆ 6 │ │ 7 ┆ 7 │ │ 8 ┆ 8 │ │ 9 ┆ 9 │ │ 10 ┆ 10 │ │ 1 ┆ 9 │ │ 2 ┆ 10 │ └───────┴───────┘
3. 用Polars内置函数简化(可选)
如果是对单行的多个条件做逻辑与/或,可直接使用Polars内置的pl.all_horizontal()或pl.any_horizontal(),代码更简洁:
# 逻辑与等价写法 df_filtered = df.filter(pl.all_horizontal(filters)) # 逻辑或等价写法 df_filtered = df.filter(pl.any_horizontal(filters))
内容的提问来源于stack exchange,提问作者Guz
相关产品推荐
相关产品推荐

