You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为新行计算特征时更新Pandas DataFrame的惯用方法

时序特征工程的Pandas高效实现方案

需求背景

拥有含5列时序数据的Pandas DataFrame,需实现compute_features函数生成20+特征列用于机器学习。实时场景下每15分钟接收一条新的5列数据行,需将其计算特征后追加到DataFrame中,其中部分特征依赖历史N个值的滚动窗口计算。要求统一特征工程逻辑,支持全量初始化计算与增量更新两种模式。

核心思路

  1. 拆分特征类型:将特征分为两类分别处理
    • 单行列内特征:仅依赖当前行原始数据,无需历史
    • 滚动窗口特征:依赖历史N行数据,复用已有计算结果避免重复运算
  2. 统一函数入口:通过update参数区分全量/增量模式,增量模式下仅计算新行的特征
  3. 复用Pandas原生API:利用rolling、concat等原生方法,保证代码简洁且符合惯用风格

实现代码

import pandas as pd
from typing import Optional

def compute_features(df: pd.DataFrame, 
                     new_row: Optional[pd.DataFrame] = None, 
                     update: bool = False) -> pd.DataFrame:
    # 计算仅依赖当前行的列内特征
    def _compute_row_features(df_subset: pd.DataFrame) -> pd.DataFrame:
        df_subset['feature_1'] = df_subset['a'] * df_subset['b']
        df_subset['feature_2'] = df_subset['c'].pow(2)
        # 追加更多列内特征...
        return df_subset
    
    # 计算依赖历史的滚动窗口特征
    def _compute_rolling_features(df_full: pd.DataFrame, is_update: bool = False) -> pd.DataFrame:
        # 定义滚动窗口配置:特征名 -> (原始列名, 窗口大小, 聚合方法)
        window_specs = {
            'feature_3': ('a', 10, 'mean'),
            'feature_4': ('b', 20, 'max'),
            'feature_5': ('c', 15, 'std')
            # 追加更多滚动特征...
        }
        
        if is_update:
            # 增量模式:仅计算最后一行的滚动特征,取最小必要历史数据
            for feat_name, (col, window, agg) in window_specs.items():
                rolling_slice = df_full[col].iloc[-window:]
                df_full.loc[df_full.index[-1], feat_name] = getattr(rolling_slice.rolling(window=window), agg)().iloc[-1]
        else:
            # 全量模式:直接计算所有行的滚动特征
            for feat_name, (col, window, agg) in window_specs.items():
                df_full[feat_name] = df_full[col].rolling(window=window).agg(agg)
        
        return df_full
    
    if update and new_row is not None:
        # 增量更新流程:追加新行 → 计算列内特征 → 计算滚动特征
        new_row = new_row.reindex(columns=df.columns[:5])  # 对齐原始数据列
        df = pd.concat([df, new_row], ignore_index=True)
        # 计算新行的列内特征,合并回原DF
        df = _compute_row_features(df.tail(1)).combine_first(df)
        df = _compute_rolling_features(df, is_update=True)
    else:
        # 全量初始化流程:先列内特征,再滚动特征
        df = _compute_row_features(df)
        df = _compute_rolling_features(df)
    
    return df

# 示例用法
if __name__ == "__main__":
    # 全量初始化
    raw_data = pd.DataFrame({
        'a': range(100),
        'b': range(100, 200),
        'c': range(200, 300),
        'd': range(300, 400),
        'e': range(400, 500)
    })
    feature_df = compute_features(raw_data)
    
    # 增量更新:模拟实时接收新行
    new_row = pd.DataFrame({
        'a': [100],
        'b': [200],
        'c': [300],
        'd': [400],
        'e': [500]
    })
    feature_df = compute_features(feature_df, new_row, update=True)

关键优势

  • 符合Pandas惯用风格:复用原生API,避免自定义复杂逻辑
  • 高效增量计算:滚动特征仅针对新行计算,无需重跑全量数据,适配实时场景
  • 逻辑统一:所有特征计算逻辑封装在同一函数内,便于维护和扩展

内容的提问来源于stack exchange,提问作者harryjulian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 11:07:08