全局转换与窗口本地转换的滚动窗口聚合性能差异探究
Polars滚动窗口聚合两种等价实现的性能差异困惑
我观察到以下两种等价的滚动窗口聚合实现存在令人困惑的性能差异。原本以为采用全局转换的版本会比针对每个窗口转换的版本快很多,但实际上后者的速度几乎是前者的两倍。第二种写法实际上更便捷,因为聚合表达式完全包含在rolling代码块中,希望有人能解释这一现象。
import numpy as np import pandas as pd import polars as pl import itertools import time from datetime import timedelta rng = np.random.default_rng() date = [t.date() for t in pd.date_range("2001-01-01", "2023-12-31", freq="B")] ids = [f"id_{i+1}" for i in range(1000)] ids, date = zip(*itertools.product(ids, date)) values = np.expm1(0.01 * rng.normal(size=len(date))) frame = pl.from_dict({"id": ids, "date": date, "value": values}).lazy() # 全局转换版本 t0 = time.time() frame.select( pl.col("id"), pl.col("date"), pl.col("value").log1p().alias("log_value") ).sort( by=["id", "date"] ).rolling( index_column="date", by="id", period=timedelta(days=30) ).agg( pl.col("log_value").sum() ).select( pl.col("id"), pl.col("date"), np.expm1(pl.col("log_value")).alias("result") ).collect() t = time.time() - t0 print(f"Time for aggregation with frame global `log1p` and `expm1` is {t=:.3f}s.") >>> Time for aggregation with frame global `log1p` and `expm1` is t=0.209s. # 窗口本地转换版本 t0 = time.time() frame.rolling( index_column="date", by="id", period=timedelta(days=30) ).agg( np.expm1(pl.col("value").log1p().sum()) ).collect() t = time.time() - t0 print(f"Time for aggregation with window local `log1p` and `expm1` is {t=:.3f}s.") >>> Time for aggregation with window local `log1p` and `expm1` is t=0.119s.
内容的提问来源于stack exchange,提问作者Benjamin Trendelkamp-Schroer
相关产品推荐
相关产品推荐

