求计算时间序列偏移量的函数,实现TAIR向tre200h0时间对齐
时间序列对齐与偏移实现方案
核心思路
通过寻找最优时间偏移量,最小化TAIR与tre200h0两条曲线的距离(如均方误差MSE),再将TAIR的时间戳按该偏移量调整,生成对齐后的DataFrame。
实现步骤与代码示例(Python)
以下基于pandas处理时间序列,scipy做优化求解,numpy计算误差:
1. 数据预处理
确保时间戳为datetime类型并排序:
import pandas as pd import numpy as np from scipy.optimize import minimize # 假设已加载数据为两个DataFrame:df_tair(含TAIR列与timestamp列)、df_tre(含tre200h0列与timestamp列) df_tair['timestamp'] = pd.to_datetime(df_tair['timestamp']) df_tre['timestamp'] = pd.to_datetime(df_tre['timestamp']) # 按时间排序,保证序列连续性 df_tair = df_tair.sort_values('timestamp').reset_index(drop=True) df_tre = df_tre.sort_values('timestamp').reset_index(drop=True)
2. 定义距离计算函数
以均方误差(MSE)为指标,计算偏移后TAIR与tre200h0的匹配误差:
def compute_offset_error(offset_hours, tair_df, tre_df): # 对TAIR时间戳施加偏移 shifted_tair = tair_df.copy() shifted_tair['timestamp'] += pd.Timedelta(hours=offset_hours) # 按最近时间戳合并两个序列,保留有效匹配 merged = pd.merge_asof(tre_df, shifted_tair, on='timestamp', direction='nearest').dropna(subset=['TAIR', 'tre200h0']) # 计算均方误差 return np.mean((merged['TAIR'] - merged['tre200h0']) ** 2)
3. 求解最优偏移量
以你目测的22小时为初始值,在合理范围内搜索最小化误差的偏移量:
# 初始猜测值与搜索范围(可根据数据调整范围) initial_guess = 22 search_bounds = [(18, 26)] # 执行最小化求解 optimization_result = minimize(compute_offset_error, initial_guess, args=(df_tair, df_tre), bounds=search_bounds) optimal_offset = optimization_result.x[0] print(f"计算得到的最优偏移量:{optimal_offset:.2f}小时")
4. 生成对齐后的TAIR DataFrame
根据最优偏移量调整时间戳,输出最终结果:
# 生成偏移后的TAIR DataFrame aligned_tair_df = df_tair.copy() aligned_tair_df['timestamp'] += pd.Timedelta(hours=optimal_offset) # 可选:与tre200h0的时间点完全对齐(按需选择) full_aligned_df = pd.merge_asof(df_tre, aligned_tair_df, on='timestamp', direction='nearest')
替代方案:交叉互相关法
如果时间序列采样频率固定,可通过交叉互相关直接找到偏移量:
from numpy import correlate # 先将两个序列插值到相同时间点(假设采样频率为1小时) common_timestamps = pd.date_range(start=min(df_tair['timestamp'].min(), df_tre['timestamp'].min()), end=max(df_tair['timestamp'].max(), df_tre['timestamp'].max()), freq='H') tair_interp = df_tair.set_index('timestamp').reindex(common_timestamps).interpolate()['TAIR'].values tre_interp = df_tre.set_index('timestamp').reindex(common_timestamps).interpolate()['tre200h0'].values # 计算交叉互相关 corr = correlate(tair_interp, tre_interp, mode='full') # 找到互相关最大值对应的偏移索引 max_corr_idx = np.argmax(corr) # 转换为小时偏移量(采样频率为1小时,所以索引差就是小时数) offset_hours = max_corr_idx - (len(tair_interp) - 1) print(f"互相关法得到的偏移量:{offset_hours}小时")
内容的提问来源于stack exchange,提问作者Cla FM
相关产品推荐
相关产品推荐

