Python中使用线性插值补全pandas时间序列15分钟间隔缺失数据
Pandas大时间序列数据集15分钟间隔线性插值高效解决方案
核心思路
依托pandas底层C实现的向量化运算能力,通过时间索引重采样+内置线性插值接口实现高效补全,全程无Python层循环,单文件处理速度比普通循环/自定义函数方案高10~100倍,适配超大体量文件及多文件批量处理需求。
单文件处理代码
1. 读取文件优化
读取阶段直接指定数据类型和时间列解析规则,避免后续类型转换开销:
import pandas as pd df = pd.read_csv( "你的源文件路径.csv", parse_dates=["Date and Time"], # 读取时直接转换为datetime类型,性能远高于后续astype转换 dtype={ "Seconds": "float64", "Pressure (mmHg)": "float32", "Temperature (C)": "float32" # 精度符合需求的前提下用float32,内存占用减半,处理速度提升明显 } )
2. 重采样+插值处理
# 设置时间为索引,若原始数据已按时间升序排列可省略sort_index进一步提速 df = df.set_index("Date and Time").sort_index() # 重采样为15分钟间隔后直接线性插值,仅对数值列生效 df_resampled = df.resample("15T").interpolate(method="linear") # 直接根据时间索引重新生成Seconds列,比插值计算更准确、速度更快 df_resampled["Seconds"] = (df_resampled.index - df_resampled.index[0]).total_seconds() # 可选:将时间索引还原为普通列 df_resampled = df_resampled.reset_index() # 保存结果 df_resampled.to_csv("处理后的文件路径.csv", index=False)
多文件批量并行处理方案
多个文件独立无依赖的场景下,可通过多进程并行进一步提升处理效率:
import pandas as pd import glob from concurrent.futures import ProcessPoolExecutor def process_single_file(file_path: str): df = pd.read_csv( file_path, parse_dates=["Date and Time"], dtype={ "Seconds": "float64", "Pressure (mmHg)": "float32", "Temperature (C)": "float32" } ) df = df.set_index("Date and Time") df_resampled = df.resample("15T").interpolate(method="linear") df_resampled["Seconds"] = (df_resampled.index - df_resampled.index[0]).total_seconds() df_resampled.reset_index().to_csv(f"processed_{file_path}", index=False) return f"{file_path} 处理完成" if __name__ == "__main__": # 匹配目标文件夹下所有csv文件,可自行修改规则 file_list = glob.glob("./目标文件夹/*.csv") # max_workers建议设置为CPU核心数-1,避免占满系统资源 with ProcessPoolExecutor(max_workers=6) as executor: for res in executor.map(process_single_file, file_list): print(res)
注意事项
- 若原始数据存在时间乱序问题,必须执行
sort_index否则插值结果会出错 - 若对压力、温度精度要求极高,可将dtype改为
float64,仅损失少量处理速度
内容的提问来源于stack exchange,提问作者Bob
相关产品推荐
相关产品推荐

