You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何精简Polars中冗长的递归表达式?

解决Polars递归表达式膨胀问题

问题背景

Polars表达式语言功能强大,但递归定义的表达式会随迭代次数增加急剧膨胀。比如示例中的diffuse函数,当n_time_steps>10时生成的表达式体积可达数GB,引发性能问题。目前通过分块调用能缓解,但希望找到更优的Polars原生方案,且要求函数输入输出保持pl.Expr类型。

现有实现代码

原始Expr级diffuse函数

import polars as pl

def diffuse(c: pl.Expr, n_time_steps=1, conductivity: float=0.05) -> pl.Expr:
    '''基于热方程的时间序列平滑函数'''
    c_conductivity = (1 - conductivity)
    new_c = c
    for _ in range(n_time_steps):
        new_c = (c_conductivity * new_c) + (0.5 * conductivity * (new_c.shift(-1).forward_fill() + new_c.shift(1).backward_fill()))
    return new_c

分块处理缓解方案

通过多次分块调用diffuse,避免单次生成超大表达式:

df = (df
       .with_columns(__temp__ = pl.col('heat'))
       .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5))  # 等效于总步数5
       .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5))  # 等效于总步数10
       .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5))  # 等效于总步数15
       .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5))  # 等效于总步数20
       .rename({'__temp__': 'heat[diffused with n_time_steps=20]'})
)

DataFrame级实现方案

通过迭代更新DataFrame列的方式,避免表达式树膨胀:

import numpy as np
import polars as pl

def add_diffused_column(
        df: pl.DataFrame, 
        col: str, 
        n_time_steps=1, 
        conductivity: float=0.05
    ) -> pl.DataFrame:
    dummy_col = f'__dummy_col_{np.random.rand()}'
    assert dummy_col not in df.columns, "随机生成的临时列名冲突"
    
    df = df.with_columns(pl.col(col).alias(dummy_col))
    for _ in range(n_time_steps):
        df = df.with_columns(pl.col(dummy_col).pipe(diffuse, n_time_steps=1, conductivity=conductivity))
    
    return df.rename({dummy_col: f'{col}[diffused with n_time_steps={n_time_steps}]'})

# 使用方式
df.pipe(add_diffused_column, 'heat', n_time_steps=100)

优化方向推荐

  • 利用Polars的fold函数优化表达式结构
    Polars的fold可在表达式层面实现迭代逻辑,且内部会优化表达式树,避免冗余膨胀。将单次迭代逻辑封装后用fold替代显式循环:

    def diffuse_iter(c: pl.Expr, conductivity: float) -> pl.Expr:
        c_conductivity = (1 - conductivity)
        return (c_conductivity * c) + (0.5 * conductivity * (c.shift(-1).forward_fill() + c.shift(1).backward_fill()))
    
    def diffuse_opt(c: pl.Expr, n_time_steps=1, conductivity: float=0.05) -> pl.Expr:
        return pl.fold(
            acc=c,
            function=lambda acc, _: diffuse_iter(acc, conductivity),
            iterable=range(n_time_steps)
        )
    

    该方案保持pl.Expr输入输出类型,同时避免生成超大表达式树。

  • 向量化运算替代循环
    若时间序列规则,可将平滑操作转化为卷积或矩阵运算,利用Polars的向量化能力一次性完成计算,彻底避免迭代式表达式构建。

  • 物化中间结果
    分块处理时,对中间结果调用.cache()物化计算结果,避免重复解析庞大的表达式树,提升后续迭代的效率:

    df = (df
           .with_columns(__temp__ = pl.col('heat'))
           .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5).cache())
           .with_columns(__temp__ = pl.col('__temp__').pipe(diffuse, n_time_steps=5).cache())
           # 继续后续分块调用
    )
    

内容的提问来源于stack exchange,提问作者MYK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 17:02:52