如何在Datetime索引DataFrame中查找前后时间戳并插值温度?
问题描述
我有一个以Datetime类型为索引的DataFrame df_temp,索引值为[2023-05-10 17:00:00, 2023-05-10 18:00:00, 2023-05-10 19:00:00],需要查找指定时间戳timestamp=2023-05-10 17:49:45.551401对应的前一个和后一个索引值(预期结果为[2023-05-10 17:00:00, 2023-05-10 18:00:00])。该DataFrame包含两种温度数据列,我需要对指定时间戳的温度进行插值计算。
尝试以下代码获取前后时间戳时触发错误:
next =df_temp.index[(timestamp-[df_temp.index[df_temp.index > timestamp]])/np.timedelta64(1,'D').TimedeltaIndex()] prev =df_temp.index[([df_temp.index[df_temp.index < timestamp]]-timestamp)/np.timedelta64(1,'D').TimedeltaIndex()]
报错信息:
TypeError: unsupported operand type(s) for -: 'Timestamp' and 'list'
同时不清楚如何在Python中对Datetime对象进行插值操作,求解决方案。
解决方案
1. 修复前后时间戳获取逻辑
报错原因是你将筛选后的索引用方括号包裹成了列表,而Timestamp和列表无法直接执行减法运算。利用Pandas DatetimeIndex的有序特性,可以通过以下两种方式快速获取目标时间的前后索引:
方法一:通过插入位置定位
import pandas as pd timestamp = pd.Timestamp('2023-05-10 17:49:45.551401') # 获取目标时间戳在索引中的插入位置(method='bfill'返回第一个大于目标值的索引位置) pos = df_temp.index.get_loc(timestamp, method='bfill') prev_idx = df_temp.index[pos-1] next_idx = df_temp.index[pos]
方法二:直接筛选极值
# 取所有小于目标时间的索引中的最大值作为前一个索引 prev_idx = df_temp.index[df_temp.index < timestamp].max() # 取所有大于目标时间的索引中的最小值作为后一个索引 next_idx = df_temp.index[df_temp.index > timestamp].min()
2. 温度插值计算
针对时间序列的插值,Pandas提供了便捷的内置方法,也可以手动实现线性插值验证:
方法一:用Pandas内置插值(推荐)
# 构造包含目标时间戳的新索引并排序 new_index = df_temp.index.append(pd.Index([timestamp])).sort_values() # 重新索引并基于时间间隔进行线性插值 interpolated_df = df_temp.reindex(new_index).interpolate(method='time') # 获取目标时间戳的插值结果 target_temperatures = interpolated_df.loc[timestamp]
method='time'会根据时间间隔自动计算线性插值,完全适配时间序列场景。
方法二:手动实现线性插值
# 获取前后索引对应的温度数据 prev_temp = df_temp.loc[prev_idx] next_temp = df_temp.loc[next_idx] # 计算时间占比 total_seconds = (next_idx - prev_idx).total_seconds() elapsed_seconds = (timestamp - prev_idx).total_seconds() ratio = elapsed_seconds / total_seconds # 计算线性插值结果 interpolated_temp = prev_temp + (next_temp - prev_temp) * ratio
内容的提问来源于stack exchange,提问作者Lisa Ulbrich
相关产品推荐
相关产品推荐

