You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求计算时间序列偏移量的函数,实现TAIR向tre200h0时间对齐

时间序列对齐与偏移实现方案

核心思路

通过寻找最优时间偏移量,最小化TAIR与tre200h0两条曲线的距离(如均方误差MSE),再将TAIR的时间戳按该偏移量调整,生成对齐后的DataFrame。


实现步骤与代码示例(Python)

以下基于pandas处理时间序列,scipy做优化求解,numpy计算误差:

1. 数据预处理

确保时间戳为datetime类型并排序:

import pandas as pd
import numpy as np
from scipy.optimize import minimize

# 假设已加载数据为两个DataFrame:df_tair(含TAIR列与timestamp列)、df_tre(含tre200h0列与timestamp列)
df_tair['timestamp'] = pd.to_datetime(df_tair['timestamp'])
df_tre['timestamp'] = pd.to_datetime(df_tre['timestamp'])

# 按时间排序,保证序列连续性
df_tair = df_tair.sort_values('timestamp').reset_index(drop=True)
df_tre = df_tre.sort_values('timestamp').reset_index(drop=True)

2. 定义距离计算函数

以均方误差(MSE)为指标,计算偏移后TAIR与tre200h0的匹配误差:

def compute_offset_error(offset_hours, tair_df, tre_df):
    # 对TAIR时间戳施加偏移
    shifted_tair = tair_df.copy()
    shifted_tair['timestamp'] += pd.Timedelta(hours=offset_hours)
    
    # 按最近时间戳合并两个序列,保留有效匹配
    merged = pd.merge_asof(tre_df, shifted_tair, on='timestamp', direction='nearest').dropna(subset=['TAIR', 'tre200h0'])
    
    # 计算均方误差
    return np.mean((merged['TAIR'] - merged['tre200h0']) ** 2)

3. 求解最优偏移量

以你目测的22小时为初始值,在合理范围内搜索最小化误差的偏移量:

# 初始猜测值与搜索范围(可根据数据调整范围)
initial_guess = 22
search_bounds = [(18, 26)]

# 执行最小化求解
optimization_result = minimize(compute_offset_error, initial_guess, args=(df_tair, df_tre), bounds=search_bounds)
optimal_offset = optimization_result.x[0]

print(f"计算得到的最优偏移量:{optimal_offset:.2f}小时")

4. 生成对齐后的TAIR DataFrame

根据最优偏移量调整时间戳,输出最终结果:

# 生成偏移后的TAIR DataFrame
aligned_tair_df = df_tair.copy()
aligned_tair_df['timestamp'] += pd.Timedelta(hours=optimal_offset)

# 可选:与tre200h0的时间点完全对齐(按需选择)
full_aligned_df = pd.merge_asof(df_tre, aligned_tair_df, on='timestamp', direction='nearest')

替代方案:交叉互相关法

如果时间序列采样频率固定,可通过交叉互相关直接找到偏移量:

from numpy import correlate

# 先将两个序列插值到相同时间点(假设采样频率为1小时)
common_timestamps = pd.date_range(start=min(df_tair['timestamp'].min(), df_tre['timestamp'].min()),
                                  end=max(df_tair['timestamp'].max(), df_tre['timestamp'].max()),
                                  freq='H')

tair_interp = df_tair.set_index('timestamp').reindex(common_timestamps).interpolate()['TAIR'].values
tre_interp = df_tre.set_index('timestamp').reindex(common_timestamps).interpolate()['tre200h0'].values

# 计算交叉互相关
corr = correlate(tair_interp, tre_interp, mode='full')
# 找到互相关最大值对应的偏移索引
max_corr_idx = np.argmax(corr)
# 转换为小时偏移量(采样频率为1小时,所以索引差就是小时数)
offset_hours = max_corr_idx - (len(tair_interp) - 1)

print(f"互相关法得到的偏移量:{offset_hours}小时")

内容的提问来源于stack exchange,提问作者Cla FM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.13 18:55:14