You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

对列表中DataFrame执行分段线性插值的技术求助

基于月份的分段线性插值批量处理方案

问题根源

你使用interpolate(method='index')未得到预期结果,是因为该方法默认基于行的位置索引计算插值,而非业务需要的month列数值。要实现分段线性插值,必须以month作为插值的基准轴。

批量处理实现

以下两种方式均可批量处理列表中的所有DataFrame,核心都是指定month为插值的参考维度:

方式1:临时设置月份为索引(兼容所有pandas版本)

遍历列表,将month设为索引后执行插值,再恢复原结构:

import pandas as pd
import numpy as np

# 示例数据初始化
d = {'ID':[5,5,5,5,5,5,5], 'month':[0,3,6,9,12,15,18], 'num':[7,np.nan,5,np.nan,np.nan,5,8]}
tempo = pd.DataFrame(data=d)
d2 = {'ID':[6,6,6,6,6,6,6], 'month':[0,3,6,9,12,15,18], 'num':[5,np.nan,2,np.nan,np.nan,np.nan,7]}
tempo2 = pd.DataFrame(data=d2)
df_list = [tempo, tempo2]

# 批量插值处理
processed_list = []
for df in df_list:
    # 以month为索引执行线性插值,再重置索引保留原列结构
    interpolated = df.set_index('month').interpolate(method='index').reset_index()
    # 还原ID列(插值不改变ID,直接从原DataFrame复制)
    interpolated['ID'] = df['ID']
    processed_list.append(interpolated)

# 打印验证结果
for idx, df in enumerate(processed_list):
    print(f"处理后的DataFrame {idx+1}:")
    print(df)
    print("---")

方式2:直接指定插值自变量(pandas 1.3.0+适用)

如果你的pandas版本≥1.3.0,可直接通过x参数指定month为插值的x轴,无需修改索引,代码更简洁:

# 批量处理一行搞定
processed_list = [df.assign(num=df['num'].interpolate(method='linear', x=df['month'])) for df in df_list]

# 验证结果
for df in processed_list:
    print(df)

结果说明

以示例中的tempo2为例,插值后num列的分段线性结果为:

  • month=3: 3.5(5与2的中点)
  • month=9: 3.666...(从month6到month15的斜率为5/9,计算得2 + (9-6)*(5/9))
  • month=12: 5.333...(2 + (12-6)*(5/9))

内容的提问来源于stack exchange,提问作者reksapj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 07:55:27