Polars滚动窗口如何跳过前N-1行?解决autocorrelation报错问题
Polars滚动窗口计算:跳过前N-1行的解决方案
问题场景
你的代码中,滚动窗口前N-1个分组(N=200)数据量不足,导致autocorrelation计算触发PanicException: Cannot apply operation on arrays of different lengths错误。要实现跳过这些窗口长度不足的行,可通过以下两种方式解决:
原始代码
import polars as pl import numpy as np import pandas as pd from functime.feature_extractors import FeatureExtractor, binned_entropy Data_Test = pl.read_csv('./DataDemo.csv') Calc_T = 200 Data_Test.with_columns( pl.col('datetime').cum_count().cast(pl.Int32).alias('Index') ).rolling('Index', period = '{}i'.format(Calc_T)).agg( pl.col('datetime').last(), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'), # pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags = 20).name.suffix('_autocorrelation') )
方法1:利用min_periods参数过滤无效窗口
通过设置rolling的min_periods为窗口长度Calc_T,强制仅当窗口内数据量≥200时才执行计算,不足的行会返回null,最后过滤掉这些无效行:
Calc_T = 200 result = ( Data_Test.with_columns( pl.col('datetime').cum_count().cast(pl.Int32).alias('Index') ) .rolling('Index', period=f'{Calc_T}i', min_periods=Calc_T) .agg( pl.col('datetime').last(), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags=20).name.suffix('_autocorrelation') ) .filter(pl.col('datetime').is_not_null()) )
方法2:直接截取起始行
从第Calc_T行开始计算滚动窗口,通过slice方法截断前N-1行数据:
Calc_T = 200 data_with_index = Data_Test.with_columns( pl.col('datetime').cum_count().cast(pl.Int32).alias('Index') ) result = ( data_with_index.slice(Calc_T - 1, None) .rolling('Index', period=f'{Calc_T}i') .agg( pl.col('datetime').last(), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.binned_entropy(bin_count=10).name.suffix('_binned_entropy'), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.absolute_energy().name.suffix('_absolute_energy'), pl.exclude('datetime', 'order_book_id', 'underlying_symbol', 'Index').ts.autocorrelation(n_lags=20).name.suffix('_autocorrelation') ) )
两种方法各有侧重:方法1保留原数据结构,仅过滤无效行;方法2直接从有效窗口起始点计算,减少不必要的计算开销,可根据实际需求选择。
内容的提问来源于stack exchange,提问作者ZaixinDong
相关产品推荐
相关产品推荐

