You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas/matplotlib新手求助:如何聚合索引不同的时间序列数据?

多索引差异时间序列的均值折线绘制方案

核心思路

通过统一x轴网格+线性插值+聚合均值的流程处理:先提取所有序列的时间边界,生成密集的统一x网格;对每个序列按该网格做线性插值,确保所有序列对齐到相同x值;最后对插值后的序列计算均值,再绘制折线。整个流程完全基于pandas实现,支持任意数量的序列(包括100+)。

具体步骤与代码示例

假设你有一个包含所有时间序列的列表all_series,每个元素是带seconds_since_start和Value列的DataFrame:

  1. 生成统一x网格
    先获取所有序列的时间边界,生成间隔均匀的x值(间隔可按需调整):

    import pandas as pd
    import matplotlib.pyplot as plt
    
    # 获取所有序列的时间范围边界
    min_x = min(df['seconds_since_start'].min() for df in all_series)
    max_x = max(df['seconds_since_start'].max() for df in all_series)
    
    # 生成间隔为0.1秒的统一x网格(可自定义间隔)
    common_x = pd.Series(pd.np.arange(min_x, max_x + 0.1, 0.1))
    common_x.name = 'seconds_since_start'
    
  2. 对每个序列做线性插值对齐
    遍历所有序列,将每个序列插值到统一x网格上:

    interpolated_dfs = []
    for df in all_series:
        # 设置时间列为索引,方便插值操作
        df_indexed = df.set_index('seconds_since_start')
        # 按统一x网格重索引并做线性插值
        interpolated = df_indexed.reindex(common_x.index).interpolate(method='linear')
        interpolated_dfs.append(interpolated)
    
  3. 合并序列并计算均值
    将所有插值结果合并,按x值计算对应均值:

    # 按列合并所有插值后的序列
    merged = pd.concat(interpolated_dfs, axis=1)
    # 计算每个x值对应的所有序列均值
    mean_series = merged.mean(axis=1).reset_index()
    mean_series.columns = ['seconds_since_start', 'Mean_Value']
    
  4. 绘制均值折线
    用matplotlib绘制最终的均值折线:

    plt.plot(mean_series['seconds_since_start'], mean_series['Mean_Value'])
    plt.xlabel('seconds_since_start')
    plt.ylabel('Mean Value')
    plt.title('Mean of Multiple Time Series')
    plt.show()
    

关键说明

  • 插值方法:interpolate(method='linear')满足你“两点间线性关系”的假设,也可根据需求替换为nearest等其他插值方式。
  • 性能适配:100+序列的场景下,pandas的concat和mean操作效率足够;若数据量极大,可考虑用dask做并行处理,常规场景下无需额外优化。
  • 网格灵活性:统一x网格的间隔可根据数据密度调整,稀疏数据用大间隔、密集数据用小间隔,平衡精度与计算量。

内容的提问来源于stack exchange,提问作者PlankTon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 07:12:39