You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars滚动窗口如何跳过前N-1行?解决autocorrelation报错问题

Polars滚动窗口计算:跳过前N-1行的解决方案

问题场景

你的代码中,滚动窗口前N-1个分组(N=200)数据量不足,导致autocorrelation计算触发PanicException: Cannot apply operation on arrays of different lengths错误。要实现跳过这些窗口长度不足的行,可通过以下两种方式解决:

原始代码

import polars as pl
import numpy as np
import pandas as pd
from functime.feature_extractors import FeatureExtractor, binned_entropy

Data_Test = pl.read_csv('./DataDemo.csv')

Calc_T = 200
Data_Test.with_columns(
    pl.col('datetime').cum_count().cast(pl.Int32).alias('Index')
).rolling('Index', period = '{}i'.format(Calc_T)).agg(
    pl.col('datetime').last(),
    pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'),
    pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'),
    # pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags = 20).name.suffix('_autocorrelation')
    
)

方法1:利用min_periods参数过滤无效窗口

通过设置rolling的min_periods为窗口长度Calc_T,强制仅当窗口内数据量≥200时才执行计算,不足的行会返回null,最后过滤掉这些无效行:

Calc_T = 200
result = (
    Data_Test.with_columns(
        pl.col('datetime').cum_count().cast(pl.Int32).alias('Index')
    )
    .rolling('Index', period=f'{Calc_T}i', min_periods=Calc_T)
    .agg(
        pl.col('datetime').last(),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags=20).name.suffix('_autocorrelation')
    )
    .filter(pl.col('datetime').is_not_null())
)

方法2:直接截取起始行

从第Calc_T行开始计算滚动窗口,通过slice方法截断前N-1行数据:

Calc_T = 200
data_with_index = Data_Test.with_columns(
    pl.col('datetime').cum_count().cast(pl.Int32).alias('Index')
)

result = (
    data_with_index.slice(Calc_T - 1, None)
    .rolling('Index', period=f'{Calc_T}i')
    .agg(
        pl.col('datetime').last(),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'),
        pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags=20).name.suffix('_autocorrelation')
    )
)

两种方法各有侧重:方法1保留原数据结构,仅过滤无效行;方法2直接从有效窗口起始点计算,减少不必要的计算开销,可根据实际需求选择。

内容的提问来源于stack exchange,提问作者ZaixinDong

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 00:32:01