You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过pd.Grouper()和.apply()获取5秒时间窗口的首尾元素?

使用pd.Grouper() + apply()获取5秒窗口的首尾元素

没问题,我懂你想坚持用pd.Grouper()配合apply()来实现这个需求,而不是切换到resample()。下面是具体的实现方案:

核心思路

当你用pd.Grouper(freq='5S')完成时间窗口分组后,每个分组都是对应窗口内的子DataFrame。只要你的原始数据是按时间戳递增排序的(从你给出的数据来看确实是这样),直接用iloc[0]和iloc[-1]就能拿到每组的首尾元素。

代码实现

方案1:仅获取首尾price

如果只需要每个窗口的首尾价格,可以写一个简洁的自定义函数:

def extract_first_last(x):
    return pd.Series({
        'first_price': x['price'].iloc[0],
        'last_price': x['price'].iloc[-1]
    })

# 分组并应用函数
result = tick.groupby(pd.Grouper(freq='5S')).apply(extract_first_last)

方案2:保留原有统计量+首尾元素

如果你还想保留原来计算的volume和num_trades,可以把逻辑整合到同一个函数里:

import numpy as np
import pandas as pd

def window_features(x):
    # 原有统计量
    volume = np.abs(x['size']).sum()
    num_trades = x['size'].count()
    # 首尾price
    first_price = x['price'].iloc[0]
    last_price = x['price'].iloc[-1]
    
    return pd.Series(
        [first_price, last_price, volume, num_trades],
        index=['first_price', 'last_price', 'volume', 'num_trades']
    )

# 执行分组计算
tick_combined = tick.groupby(pd.Grouper(freq='5S')).apply(window_features)

注意事项

  • 确保你的DataFrame索引是datetime类型:如果之前没设置,先执行tick.index = pd.to_datetime(tick.index),否则pd.Grouper无法正确识别时间窗口。
  • 如果你的数据不是按时间排序的,建议先排序:tick = tick.sort_index(),否则iloc[0]和iloc[-1]可能不是时间上的首尾元素。

内容的提问来源于stack exchange,提问作者swifty

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:12:05