对列表中DataFrame执行分段线性插值的技术求助
基于月份的分段线性插值批量处理方案
问题根源
你使用interpolate(method='index')未得到预期结果,是因为该方法默认基于行的位置索引计算插值,而非业务需要的month列数值。要实现分段线性插值,必须以month作为插值的基准轴。
批量处理实现
以下两种方式均可批量处理列表中的所有DataFrame,核心都是指定month为插值的参考维度:
方式1:临时设置月份为索引(兼容所有pandas版本)
遍历列表,将month设为索引后执行插值,再恢复原结构:
import pandas as pd import numpy as np # 示例数据初始化 d = {'ID':[5,5,5,5,5,5,5], 'month':[0,3,6,9,12,15,18], 'num':[7,np.nan,5,np.nan,np.nan,5,8]} tempo = pd.DataFrame(data=d) d2 = {'ID':[6,6,6,6,6,6,6], 'month':[0,3,6,9,12,15,18], 'num':[5,np.nan,2,np.nan,np.nan,np.nan,7]} tempo2 = pd.DataFrame(data=d2) df_list = [tempo, tempo2] # 批量插值处理 processed_list = [] for df in df_list: # 以month为索引执行线性插值,再重置索引保留原列结构 interpolated = df.set_index('month').interpolate(method='index').reset_index() # 还原ID列(插值不改变ID,直接从原DataFrame复制) interpolated['ID'] = df['ID'] processed_list.append(interpolated) # 打印验证结果 for idx, df in enumerate(processed_list): print(f"处理后的DataFrame {idx+1}:") print(df) print("---")
方式2:直接指定插值自变量(pandas 1.3.0+适用)
如果你的pandas版本≥1.3.0,可直接通过x参数指定month为插值的x轴,无需修改索引,代码更简洁:
# 批量处理一行搞定 processed_list = [df.assign(num=df['num'].interpolate(method='linear', x=df['month'])) for df in df_list] # 验证结果 for df in processed_list: print(df)
结果说明
以示例中的tempo2为例,插值后num列的分段线性结果为:
- month=3: 3.5(5与2的中点)
- month=9: 3.666...(从month6到month15的斜率为5/9,计算得2 + (9-6)*(5/9))
- month=12: 5.333...(2 + (12-6)*(5/9))
内容的提问来源于stack exchange,提问作者reksapj
相关产品推荐
相关产品推荐

