如何将泊松CDF编写为Python Polars惰性表达式?
用Polars惰性表达式实现泊松CDF特征
Polars目前没有内置的泊松CDF表达式,但可以通过map_batches封装scipy的向量化计算,既保留惰性执行的优势,又避免逐行apply的性能瓶颈。
推荐方案:封装自定义惰性表达式
1. 编写表达式函数
这个函数返回Polars表达式,利用map_batches批量处理整列数据,全程支持惰性执行:
import polars as pl from scipy.stats import poisson def poisson_cdf(count_col: str = "count", expected_col: str = "expected_count", alias: str = "poisson_cdf") -> pl.Expr: return pl.struct([count_col, expected_col]).map_batches( lambda batch: poisson.cdf( batch.struct.field(count_col).to_numpy(), batch.struct.field(expected_col).to_numpy() ), return_dtype=pl.Float64 ).alias(alias)
2. 在惰性管道中使用
不管是普通DataFrame还是LazyFrame,都可以直接嵌入到查询计划中:
普通DataFrame示例
df = pl.DataFrame({ "count": [9,2,3,4,5], "expected_count": [7.7, 0.2, 0.7, 1.1, 7.5] }) # 直接在select/with_columns中调用 result_df = df.select( pl.all(), # 保留原有列 poisson_cdf() ) print(result_df)
LazyFrame惰性执行示例
lazy_df = pl.LazyFrame({ "count": [9,2,3,4,5], "expected_count": [7.7, 0.2, 0.7, 1.1, 7.5] }) # 构建惰性查询计划,collect时才执行计算 result_df = lazy_df.select( pl.all(), poisson_cdf() ).collect() print(result_df)
性能说明
map_batches是批量向量化处理:将整列数据转为numpy数组后传给scipy,scipy内部用C实现的优化逻辑计算,性能远高于逐行apply,适合百万级以上数据集。- 完全兼容惰性执行:表达式会被纳入Polars的查询优化计划,享受缓存、谓词下推等惰性执行的优势。
备选方案:不依赖scipy核心库的实现
如果不想引入scipy.stats依赖,可以用scipy.special的伽马函数手动实现泊松CDF(泊松CDF等价于1减去上不完全伽马函数):
import polars as pl from scipy.special import gammaincc def poisson_cdf_light(count_col: str = "count", expected_col: str = "expected_count", alias: str = "poisson_cdf") -> pl.Expr: return pl.struct([count_col, expected_col]).map_batches( lambda batch: 1 - gammaincc( batch.struct.field(count_col) + 1, batch.struct.field(expected_col) ), return_dtype=pl.Float64 ).alias(alias)
若要完全脱离scipy依赖,需要自行实现阶乘与指数项的向量化计算,但精度和效率会不如scipy的优化版本,仅推荐在特殊场景下使用。
内容的提问来源于stack exchange,提问作者TheRealBenbo
相关产品推荐
相关产品推荐

