如何下采样周期时间序列并保留模式以实现周期性行为特征分析?
下采样时间序列的周期性/近周期性检测最佳实践
一、合成数据生成
我正在研究下采样后时间序列数据的特征分析,参考资料后生成了粒度为5分钟的周期合成数据(每小时12条记录,全天共288条),支持PWM信号、正三角脉冲、矩形正脉冲、冲激等模式。生成PWM调制正弦信号的代码如下:
import numpy as np import pandas as pd import scipy.signal as signal import matplotlib.pyplot as plt # Generate periodic data : PWM Modulated Sinusoidal Signal # Set the duty cycle percentage for PWM percent = 40.0 # Set the time period for one cycle of the PWM signal TimePeriod = 6.0 # Set the desired number of samples desired_samples = 278 # Calculate the time step (dt) to achieve the desired number of samples # 30 is the original number of cycles in the provided code dt = TimePeriod / (desired_samples / 30) # Calculate the number of cycles needed to achieve the desired number of samples Cycles = int(desired_samples * dt / TimePeriod) + 1 # Create a time array t = np.arange(0, Cycles * TimePeriod, dt) # Generate a PWM signal pwm = (t % TimePeriod) < (TimePeriod * percent / 100) # Create a sinusoidal signal x = np.linspace(-10, 10, len(pwm)) y_pwm = np.sin(x) # Zero out the sinusoidal signal where PWM is zero y_pwm[pwm == 0] = 0 # Convert data to a Pandas DataFrame data = {'datetime': t_num, "PWM":y_pwm} df = pd.DataFrame(data) df.shape #(288, 2)
二、下采样处理流程
得到含datetime时间戳和PWM信号的单变量序列后,先通过pd.to_timedelta和pd.to_datetime格式化时间列,再用两种方法将5分钟粒度下采样为1小时粒度(采用mean()聚合):
方法1:使用resample()
resampled_df = (df.set_index('datetime') # 将时间列设为索引 .resample('1H') # 按1小时粒度重采样 .mean() # 均值聚合 .interpolate() # 填充缺失值(可选) ) resampled_df.shape # (24, 1)
方法2:使用groupby+pd.Grouper
resampled_df2 = (df.set_index('datetime') # 将时间列设为索引 .groupby([pd.Grouper(freq='1H')]) # 按1小时分组 .mean() # 均值聚合 ) resampled_df2.shape # (24, 1)
处理后已通过绘图对比原始数据与下采样数据。
三、核心需求与疑问
目标是检测时间序列的周期性/循环行为以开展特征分析,但因数据量要求必须下采样,且不同分辨率可能存在近周期性模式。现寻求下采样+周期性/近周期性检测的最佳实践,同时有以下限制:
- 不希望使用
.iloc按位置选取数据的方法 - 已知下采样会丢失信息
- 对FFT/DFT是否适用于该场景存疑
最佳实践建议
1. 先检测原始数据周期,再针对性下采样
下采样的核心风险是破坏原始周期特征,建议先在原始5分钟粒度数据上完成周期性检测,再根据结果选择下采样策略:
- FFT/DFT完全适用:对周期性、近周期性信号的频率提取效果稳定。可通过
scipy.fft.fft计算频谱,找到峰值对应的频率(周期=1/频率)。注意先通过scipy.signal.detrend去除信号趋势,避免直流分量干扰。 - 自相关函数(ACF):计算序列自相关系数,峰值对应的滞后步长即为周期。用
pd.plotting.autocorrelation_plot快速可视化,statsmodels.tsa.stattools.acf计算具体数值。
得到原始周期后,选择下采样频率时需确保下采样间隔是原始周期的整数约数/倍数,减少信息丢失。比如原始周期为6小时,下采样为1小时或2小时粒度,能更好保留周期特征。
2. 下采样时选择适配信号类型的聚合方式
不要局限于mean(),根据信号特性调整:
- 脉冲类信号(矩形脉冲、冲激):用
max()或sum()保留脉冲强度,mean()会稀释脉冲特征 - 连续周期信号(PWM、三角波):
mean()或median()均可,若信号有明显峰值,可同时保留mean()和max()作为特征
3. 多分辨率周期检测
若不同分辨率存在不同近周期模式,采用多尺度分析:
- 小波变换(如
pywt库):能在不同时间尺度上检测周期,同时保留时间局部性,适合近周期性信号 - 对不同下采样粒度(如1小时、2小时、6小时)分别做周期检测,对比结果提取多尺度周期特征
4. 替代按位置采样的方案
完全基于时间索引操作,避免.iloc:
- 沿用你已使用的
resample()或groupby+pd.Grouper,二者均基于时间逻辑采样,而非位置 - 若需自定义采样窗口,可使用
pd.Grouper的offset参数调整窗口起始时间,或用rolling()结合时间窗口做滑动聚合
内容的提问来源于stack exchange,提问作者Mario
相关产品推荐
相关产品推荐

