多列CSV数据1000时间步插值需求(无Pandas/循环)
多列NumPy数组的时间序列插值(无Pandas/循环)
问题描述
我有一份以NumPy数组存储的多列数据,原始格式如下:
[( 7, 1818, 1, 8, 1818.021, 65, 10.2, 1, 1) ( 12, 1818, 1, 13, 1818.034, 37, 7.7, 1, 1) ( 16, 1818, 1, 17, 1818.045, 77, 11.1, 1, 1) ... (73715, 2019, 10, 29, 2019.826, 0, 0. , 30, 0) (73716, 2019, 10, 30, 2019.829, 0, 0. , 24, 0) (73717, 2019, 10, 31, 2019.832, 0, 0. , 28, 0)]
其中第2列为年份。需要将1900至2000年的数据插值为1000个时间步,且每列数据需独立完成插值。
已通过以下代码提取目标时间段数据(共36889条):
index1 = np.where(arr['year']==1900)[0][0] index2 = np.where(arr['year']==2000)[0][-1] data = arr[index1:index2]
提取后数据格式:
[(29950, 1900, 1, 1, 1900.001, 12, 3. , 1, 1) (29951, 1900, 1, 2, 1900.004, 12, 3. , 1, 1) (29952, 1900, 1, 3, 1900.007, 3, 2. , 1, 1) ... (66836, 2000, 12, 28, 2000.99 , 162, 7. , 13, 1) (66837, 2000, 12, 29, 2000.993, 151, 11.7, 15, 1) (66838, 2000, 12, 30, 2000.996, 152, 10.6, 11, 1)]
尝试过np.interp()和scipy.interp1d但未找到多列批量插值的明确方法,现寻求不使用Pandas和显式循环的实现方案。
解决方案
利用scipy.interpolate.interp1d的多维数组支持,通过指定axis=0实现对每列的批量向量化插值,无需逐列循环。
实现步骤
- 转换数据格式:将结构化NumPy数组转为普通二维数组,便于插值函数处理。
- 选择时间基准:用第5列的连续时间值(如
1900.001)作为插值的x轴,避免离散日期的处理复杂度。 - 生成目标时间步:创建1000个均匀分布的时间点,覆盖1900到2000的完整范围。
- 批量插值:调用
interp1d并指定axis=0,对每列数据独立完成插值。
代码示例
import numpy as np from scipy.interpolate import interp1d # 将结构化数组转换为普通二维浮点数组 data_array = data.view(np.float64).reshape(data.shape + (-1,)) # 提取原始连续时间轴(第5列,索引为4) x_original = data_array[:, 4] # 生成1000个均匀分布的目标时间点 x_new = np.linspace(x_original.min(), x_original.max(), 1000) # 创建插值函数,axis=0表示对每列(行维度)插值 interp_func = interp1d(x_original, data_array, axis=0, kind='linear') # 执行插值得到结果 interpolated_data = interp_func(x_new) # 可选:将结果还原为原始结构化数组格式 interpolated_struct = interpolated_data.view(data.dtype)
关键细节
kind='linear'可替换为'cubic'、'nearest'等插值类型,根据数据特性调整。- 向量化操作完全依赖NumPy/SciPy的底层优化,效率远高于显式循环。
- 若原始数据为普通二维数组,可跳过结构化数组的转换步骤。
内容的提问来源于stack exchange,提问作者Andromeda
相关产品推荐
相关产品推荐

