如何在Polars中实现按日期计算两种特征偏度并填充列
问题:Pandas累积偏度计算迁移至Polars的实现方案
需要将计算数据框两种特征偏度的Pandas函数迁移到Polars:
- 单日偏度:
date == d时对应特征的偏度 - 累积偏度:
date <= d时对应特征的偏度
已完成单日偏度的转换,但累积偏度的实现遇到困难,以下是相关代码:
数据框生成代码
import random import pandas as pd date = [random.choice([1,2,3]) for x in range(0,100)] feature1 = [random.gauss(0,1) for x in range(0,100)] feature2 = [random.gauss(0,2) for x in range(0,100)] df = pd.DataFrame({"date":date,'feature1':feature1,'feature2':feature2})
Pandas原函数
import itertools from scipy.stats import skew def features_augment(df): dff = df.copy() for col,d in itertools.product(dff.columns[2:],dff.date.unique()): dff.loc[dff.date==d,'sk_'+col] = skew(dff.loc[dff.date==d,col]) # 单日偏度 dff.loc[dff.date==d,'rsk_'+col] = skew(dff.loc[dff.date<=d,col]) # 累积偏度 return dff
Polars未完成函数
import polars as pl from scipy.stats import skew def pl_feature_augment(df): pl_df = pl.from_pandas(df) sk = pl_df.groupby("date").agg(pl.all().exclude("id","volvol").skew().prefix("sk_")) # 计算单日特征偏度 pl_df = pl_df.join(sk,"date") for d in pl_df.select(pl.col("date")).unique(): pl_df.with_columns(pl.when(pl.col("date")<=d).then(pl.all().exclude("id","vol","volvol").skew()).prefix("rsk_")) # 无效 return pl_df.to_pandas()
解决方案
Polars中实现累积偏度不需要低效循环,提供两种可行方案:
方案1:循环处理唯一日期(逻辑清晰,适合小数据集)
import polars as pl from scipy.stats import skew def pl_feature_augment(df): pl_df = pl.from_pandas(df) feature_cols = pl_df.columns[1:] # 提取特征列(排除date) # 计算单日偏度 sk_df = pl_df.groupby("date").agg( [pl.col(col).skew().alias(f"sk_{col}") for col in feature_cols] ) pl_df = pl_df.join(sk_df, on="date") # 计算累积偏度 unique_dates = sorted(pl_df.select("date").unique().to_series().to_list()) cumulative_skew_rows = [] for d in unique_dates: # 筛选当前及之前所有日期的数据 cumulative_data = pl_df.filter(pl.col("date") <= d) # 计算每个特征的累积偏度 skew_results = cumulative_data.select( [pl.col(col).skew().alias(f"rsk_{col}") for col in feature_cols] ).row(0) # 组装行数据 row = {"date": d} row.update(dict(zip([f"rsk_{col}" for col in feature_cols], skew_results))) cumulative_skew_rows.append(row) # 转换为Polars DataFrame并关联回原数据 rsk_df = pl.DataFrame(cumulative_skew_rows) pl_df = pl_df.join(rsk_df, on="date") return pl_df.to_pandas()
方案2:纯Polars向量操作(无Python循环,适合大数据集)
利用Polars的列表聚合和累积操作,避免Python循环,提升效率:
import polars as pl def pl_feature_augment(df): pl_df = pl.from_pandas(df).sort("date") feature_cols = pl_df.columns[1:] # 计算单日偏度 sk_df = pl_df.groupby("date").agg( [pl.col(col).skew().alias(f"sk_{col}") for col in feature_cols] ) # 计算累积偏度:先按date聚合特征值列表,再计算累积合并列表的偏度 cumulative_rsk_df = ( pl_df.groupby("date") .agg([pl.col(col).alias(col) for col in feature_cols]) .sort("date") .with_columns( [ # 累积合并所有历史日期的特征值列表,再计算偏度 pl.col(col).cum_eval(lambda x: x.list.concat()).list.skew().alias(f"rsk_{col}") for col in feature_cols ] ) .select(["date"] + [f"rsk_{col}" for col in feature_cols]) ) # 合并单日偏度和累积偏度结果 pl_df = pl_df.join(sk_df, on="date").join(cumulative_rsk_df, on="date") return pl_df.to_pandas()
内容的提问来源于stack exchange,提问作者Mayeul sgc
相关产品推荐
相关产品推荐

