如何用Pandas将小时级温度数据扩展为15分钟间隔并插值?
将小时级温度数据转换为15分钟间隔的线性插值方法
需求说明
需要把小时级时间间隔的温度数据集转换为15分钟(一刻钟)间隔,中间时刻的温度按相邻小时温度变化的25%、50%、75%递增估算,示例如下:
原数据:
date_time Temperature [°C] 2018-01-01 01:00:00 10 2018-01-01 02:00:00 12 期望输出:
date_time Temperature [°C] 2018-01-01 01:00:00 10 2018-01-01 01:15:00 10.5 2018-01-01 01:30:00 11 2018-01-01 01:45:00 11.5 2018-01-01 02:00:00 12
问题:手动插入行不适用于大规模数据
最初尝试手动插入行的方式,代码如下:
df_test.loc[1.5] = ['time', 'temp_old', 'temp_new'] df_test = df_test.sort_index().reset_index(drop=True)
但该方法效率极低且易出错,无法适配5000+行的大规模数据集。
解决方案:使用Pandas的resample + interpolate
借助Pandas的resample和interpolate方法,可以高效实现需求,代码如下:
# 复制原数据集(避免修改原始数据) df_weather_test = df_weather.copy() # 将时间列转换为datetime类型(resample的必要前提) df_weather_test['date_time'] = pd.to_datetime(df_weather['date_time']) # 重采样为15分钟间隔,并用线性插值填充中间值 df_weather_test2 = df_weather_test.resample('15T', on='date_time').mean().interpolate()
代码说明
pd.to_datetime():确保时间列是标准的datetime格式,这是时间序列重采样的基础。resample('15T'):按15分钟间隔对时间序列拆分,'15T'是Pandas中代表15分钟的缩写。.mean():重采样时对每个时间点取均值(原数据为小时级,所以原时间点的均值就是自身数值)。.interpolate():默认采用线性插值,正好匹配相邻小时温度按25%、50%、75%递增的估算逻辑。
内容的提问来源于stack exchange,提问作者H.E.
相关产品推荐
相关产品推荐

