如何精简Polars中冗长的递归表达式?
解决Polars递归表达式膨胀问题
问题背景
Polars表达式语言功能强大,但递归定义的表达式会随迭代次数增加急剧膨胀。比如示例中的diffuse函数,当n_time_steps>10时生成的表达式体积可达数GB,引发性能问题。目前通过分块调用能缓解,但希望找到更优的Polars原生方案,且要求函数输入输出保持pl.Expr类型。
现有实现代码
原始Expr级diffuse函数
import polars as pl def diffuse(c: pl.Expr, n_time_steps=1, conductivity: float=0.05) -> pl.Expr: '''基于热方程的时间序列平滑函数''' c_conductivity = (1 - conductivity) new_c = c for _ in range(n_time_steps): new_c = (c_conductivity * new_c) + (0.5 * conductivity * (new_c.shift(-1).forward_fill() + new_c.shift(1).backward_fill())) return new_c
分块处理缓解方案
通过多次分块调用diffuse,避免单次生成超大表达式:
df = (df .with_columns(__temp__ = pl.col('heat')) .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5)) # 等效于总步数5 .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5)) # 等效于总步数10 .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5)) # 等效于总步数15 .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5)) # 等效于总步数20 .rename({'__temp__': 'heat[diffused with n_time_steps=20]'}) )
DataFrame级实现方案
通过迭代更新DataFrame列的方式,避免表达式树膨胀:
import numpy as np import polars as pl def add_diffused_column( df: pl.DataFrame, col: str, n_time_steps=1, conductivity: float=0.05 ) -> pl.DataFrame: dummy_col = f'__dummy_col_{np.random.rand()}' assert dummy_col not in df.columns, "随机生成的临时列名冲突" df = df.with_columns(pl.col(col).alias(dummy_col)) for _ in range(n_time_steps): df = df.with_columns(pl.col(dummy_col).pipe(diffuse, n_time_steps=1, conductivity=conductivity)) return df.rename({dummy_col: f'{col}[diffused with n_time_steps={n_time_steps}]'}) # 使用方式 df.pipe(add_diffused_column, 'heat', n_time_steps=100)
优化方向推荐
利用Polars的
fold函数优化表达式结构
Polars的fold可在表达式层面实现迭代逻辑,且内部会优化表达式树,避免冗余膨胀。将单次迭代逻辑封装后用fold替代显式循环:def diffuse_iter(c: pl.Expr, conductivity: float) -> pl.Expr: c_conductivity = (1 - conductivity) return (c_conductivity * c) + (0.5 * conductivity * (c.shift(-1).forward_fill() + c.shift(1).backward_fill())) def diffuse_opt(c: pl.Expr, n_time_steps=1, conductivity: float=0.05) -> pl.Expr: return pl.fold( acc=c, function=lambda acc, _: diffuse_iter(acc, conductivity), iterable=range(n_time_steps) )该方案保持
pl.Expr输入输出类型,同时避免生成超大表达式树。向量化运算替代循环
若时间序列规则,可将平滑操作转化为卷积或矩阵运算,利用Polars的向量化能力一次性完成计算,彻底避免迭代式表达式构建。物化中间结果
分块处理时,对中间结果调用.cache()物化计算结果,避免重复解析庞大的表达式树,提升后续迭代的效率:df = (df .with_columns(__temp__ = pl.col('heat')) .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5).cache()) .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5).cache()) # 继续后续分块调用 )
内容的提问来源于stack exchange,提问作者MYK
相关产品推荐
相关产品推荐

