You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas时间序列:聚合与重采样的差异及低信息损失降采样方法

时间序列降采样与聚合方法分析

我正在开展时间序列数据的特征分析实验,生成了如下三角脉冲时间序列数据:

import numpy as np
import pandas as pd
import scipy.signal as signal
import matplotlib.pyplot as plt

# 生成三角脉冲的函数
def generate_triangular_pulse(rise_duration, fall_duration, period, samples):
    control_points_x = np.array([0, rise_duration, rise_duration + fall_duration, period])
    control_points_y = np.array([0, 1, 0, 0])

    x = np.linspace(0, period, samples)
    return np.interp(x, control_points_x, control_points_y)

# 设置参数
samples = 288
rise_duration = 1
fall_duration = 3
period = 5
num_triangles = 5

# 创建5分钟间隔的时间序列
t_num = pd.date_range(start='2024-01-01', freq='5T', periods=samples)

# 生成三角脉冲数据
triangular_pulse = np.zeros(samples)
for i in range(num_triangles):
    start_index = i * (samples // num_triangles)
    end_index = (i + 1) * (samples // num_triangles)
    triangular_pulse[start_index:end_index] = generate_triangular_pulse(rise_duration, fall_duration, period, end_index - start_index)

# 转换为DataFrame
data = {'date': t_num, 'Positive Triangle Pulse': triangular_pulse }
df = pd.DataFrame(data)
print(df.head())

#               date  Positive Triangle Pulse
#0 2024-01-01 00:00:00                 0.000000
#1 2024-01-01 00:05:00                 0.089286
#2 2024-01-01 00:10:00                 0.178571
#3 2024-01-01 00:15:00                 0.267857
#4 2024-01-01 00:20:00                 0.357143
#288 rows × 2 columns

我希望将数据从5分钟频率降采样至1小时频率,同时尽可能减少信息损失,当前使用了如下代码:

resampled_df = (df.set_index('date')          # 修正:原数据列名为'date'而非'datetime'
                  .resample('1H')                 # 按1小时频率重采样
                  .mean()                         # 用均值聚合
                  .interpolate()                  # 填充缺失值(以防万一)
                )
resampled_df.shape                                # (24, 1)

并绘制了原始数据与降采样后数据的对比图:

import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import pandas as pd

fig, axes = plt.subplots(nrows=1, ncols=2, figsize=(15, 4))

# 绘制原始数据
axes[0].plot(df['date'], df['Positive Triangle Pulse'], "b.-", label="原始数据")
axes[0].set_title(f'正三角脉冲序列(共{len(df)}条观测)')

# 绘制降采样后数据
axes[1].plot(resampled_df.index, resampled_df['Positive Triangle Pulse'], "b.-", label="降采样数据")
axes[1].set_title(f'正三角脉冲序列(1小时频率降采样,共{len(resampled_df)}条观测)')

step_size = 12
selected_ticks = df['date'][::step_size]

for ax in axes:
    ax.set_xticks(selected_ticks)
    ax.set_xticklabels(selected_ticks, rotation=90)

plt.legend(loc="best")
plt.show()

原始与降采样数据对比图


技术问题解答

1. Pandas中resample()与agg()/groupby()聚合的差异,及记录选择/整合的区分

  • 核心定位差异

    • resample()是时间序列专用工具:基于指定时间频率自动划分窗口,自带时间对齐、缺失时间段补空等逻辑,无需手动定义分组规则,完全适配时间序列场景。
    • groupby()是通用分组工具:可基于任意列或自定义键分组,若用于时间序列,需手动将时间转换为分组键(比如提取小时、日期),本质是把时间当作普通分类变量处理,没有时间特有的逻辑支持。
    • agg()是聚合执行器:本身不做分组,需配合resample()或groupby()使用,用来指定具体聚合规则(如mean、max)。
  • 记录选择/剔除 vs 整合消化

    • 记录选择/剔除:直接从原始数据中挑选或丢弃部分记录,不生成新值。比如resample('1H').first()、groupby(小时).last(),都是保留窗口内某条原始记录,剔除其他。
    • 整合消化:对窗口内所有观测值计算生成新值,结果不是原始数据中的任何一条。比如resample('1H').mean()、groupby().sum(),属于把多个值整合为一个新值的处理。agg()只是执行这类整合的工具,具体是选择还是整合取决于搭配的聚合函数。

2. 哪种方法对时间数据模式行为影响最小?

要最小化对时间模式(比如三角脉冲的上升/下降趋势、峰值)的影响,关键是选对聚合逻辑,而非单纯看工具:

  • 若想保留脉冲峰值特征:用resample('1H').max()(或groupby按小时分组取max),能抓住每个小时窗口内的最高值,不会像均值那样抹平峰值。
  • 若想保留脉冲趋势变化:可以用resample('1H').apply()自定义聚合逻辑(比如提取窗口内的斜率、保留关键时间点的值);或使用插值类降采样(如resample('1H').interpolate(method='linear')),适合连续平滑的趋势场景。
  • 工具对比:resample()比手动groupby()更可靠,它自动处理时间窗口边界对齐,避免手动分组时的时间错位问题,能更准确保留原始模式。

对于你的三角脉冲数据,用resample配合max()或自定义关键特征的聚合函数,对原始模式的影响最小,能保留脉冲的峰值和大致形状;而均值聚合会明显抹平脉冲的尖锐特征,从对比图也能看出降采样后的曲线平缓了很多。


内容的提问来源于stack exchange,提问作者Mario

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 18:19:55