为新行计算特征时更新Pandas DataFrame的惯用方法
时序特征工程的Pandas高效实现方案
需求背景
拥有含5列时序数据的Pandas DataFrame,需实现compute_features函数生成20+特征列用于机器学习。实时场景下每15分钟接收一条新的5列数据行,需将其计算特征后追加到DataFrame中,其中部分特征依赖历史N个值的滚动窗口计算。要求统一特征工程逻辑,支持全量初始化计算与增量更新两种模式。
核心思路
- 拆分特征类型:将特征分为两类分别处理
- 单行列内特征:仅依赖当前行原始数据,无需历史
- 滚动窗口特征:依赖历史N行数据,复用已有计算结果避免重复运算
- 统一函数入口:通过
update参数区分全量/增量模式,增量模式下仅计算新行的特征 - 复用Pandas原生API:利用
rolling、concat等原生方法,保证代码简洁且符合惯用风格
实现代码
import pandas as pd from typing import Optional def compute_features(df: pd.DataFrame, new_row: Optional[pd.DataFrame] = None, update: bool = False) -> pd.DataFrame: # 计算仅依赖当前行的列内特征 def _compute_row_features(df_subset: pd.DataFrame) -> pd.DataFrame: df_subset['feature_1'] = df_subset['a'] * df_subset['b'] df_subset['feature_2'] = df_subset['c'].pow(2) # 追加更多列内特征... return df_subset # 计算依赖历史的滚动窗口特征 def _compute_rolling_features(df_full: pd.DataFrame, is_update: bool = False) -> pd.DataFrame: # 定义滚动窗口配置:特征名 -> (原始列名, 窗口大小, 聚合方法) window_specs = { 'feature_3': ('a', 10, 'mean'), 'feature_4': ('b', 20, 'max'), 'feature_5': ('c', 15, 'std') # 追加更多滚动特征... } if is_update: # 增量模式:仅计算最后一行的滚动特征,取最小必要历史数据 for feat_name, (col, window, agg) in window_specs.items(): rolling_slice = df_full[col].iloc[-window:] df_full.loc[df_full.index[-1], feat_name] = getattr(rolling_slice.rolling(window=window), agg)().iloc[-1] else: # 全量模式:直接计算所有行的滚动特征 for feat_name, (col, window, agg) in window_specs.items(): df_full[feat_name] = df_full[col].rolling(window=window).agg(agg) return df_full if update and new_row is not None: # 增量更新流程:追加新行 → 计算列内特征 → 计算滚动特征 new_row = new_row.reindex(columns=df.columns[:5]) # 对齐原始数据列 df = pd.concat([df, new_row], ignore_index=True) # 计算新行的列内特征,合并回原DF df = _compute_row_features(df.tail(1)).combine_first(df) df = _compute_rolling_features(df, is_update=True) else: # 全量初始化流程:先列内特征,再滚动特征 df = _compute_row_features(df) df = _compute_rolling_features(df) return df # 示例用法 if __name__ == "__main__": # 全量初始化 raw_data = pd.DataFrame({ 'a': range(100), 'b': range(100, 200), 'c': range(200, 300), 'd': range(300, 400), 'e': range(400, 500) }) feature_df = compute_features(raw_data) # 增量更新:模拟实时接收新行 new_row = pd.DataFrame({ 'a': [100], 'b': [200], 'c': [300], 'd': [400], 'e': [500] }) feature_df = compute_features(feature_df, new_row, update=True)
关键优势
- 符合Pandas惯用风格:复用原生API,避免自定义复杂逻辑
- 高效增量计算:滚动特征仅针对新行计算,无需重跑全量数据,适配实时场景
- 逻辑统一:所有特征计算逻辑封装在同一函数内,便于维护和扩展
内容的提问来源于stack exchange,提问作者harryjulian
相关产品推荐
相关产品推荐

