如何将不同长度的多个DataFrame绘制到同一张图表中
问题描述
我有大约300个DataFrame,希望将它们绘制到同一张折线图中。问题在于这些DataFrame的数值数量(即长度)各不相同。如何绘制才能让它们在图表的x轴上首尾对齐?
我可以将DataFrame转换为列表、Series或其他格式。
示例
- df1:
[5, 3, 10, 7, 10...] - df2:
[2, 4, 5, 7, 2, 1, 3, 0, 1] - 调整后的df2:理想状态——与df1长度相同
df1 df2 adjusted df2 5 2 2 3 2 10 2 7 4 4 10 4 5 4 8 6 6 6 6 5 6 6 7 7 2 7 1 7 5 2 2 3 2 6 2 9 1 1 9 1 7 1 10 3 3 2 3 7 3 7 0 0 6 0 1 0 6 1 1 9 1
含日期时间的示例
第一个DataFrame:
datetime value 5448 2020-01-19 22:05:00 166.300003 5449 2020-01-19 22:10:00 165.259995 5450 2020-01-19 22:15:00 164.699997 5451 2020-01-19 22:20:00 165.380005 5452 2020-01-19 22:25:00 166.179993 5453 2020-01-19 22:30:00 162.630005 5424 2020-01-19 22:35:00 162.550003 5425 2020-01-19 22:40:00 161.990005 5426 2020-01-19 22:45:00 161.750000 5427 2020-01-10 22:50:00 161.440002
第二个DataFrame:
datetime value 15900 2020-02-25 11:55:00 262.510010 15901 2020-02-25 12:00:00 263.179993 15902 2020-02-25 12:05:00 262.260010 15903 2020-02-25 12:10:00 261.959991 15904 2020-02-25 12:15:00 262.179993 15905 2020-02-25 12:20:00 261.299988 15906 2020-02-25 12:25:00 261.579987 15907 2020-02-25 12:30:00 261.890015 15908 2020-02-25 12:35:00 262.820007 15909 2020-02-25 12:40:00 262.010010 15910 2020-02-25 12:45:00 261.630005 15911 2020-02-25 12:50:00 261.109985 15912 2020-02-25 12:55:00 261.149994 15913 2020-02-25 13:00:00 260.679993 15914 2020-02-25 13:05:00 261.929993 15915 2020-02-25 13:10:00 260.880005 15916 2020-02-25 13:15:00 259.929993
第三个DataFrame:
datetime value 16407 2020-02-27 06:10:00 224.860001 16408 2020-02-27 06:15:00 224.240005 16409 2020-02-27 06:20:00 223.610001 16410 2020-02-27 06:25:00 223.490005 16411 2020-02-27 06:30:00 223.199997
期望绘制效果:所有折线左对齐(起点重合)、右对齐(终点重合)。
解决方案
方案1:使用相对位置x轴(通用方法)
不需要修改原始数据,直接将每个序列的x轴映射为0到1之间的相对索引,即每个点的x值为当前索引 / (序列长度-1)。这样不管序列长短,都会从x=0开始,x=1结束,自然首尾对齐。
示例代码(使用matplotlib):
import matplotlib.pyplot as plt import pandas as pd # 假设有多个DataFrame存储在列表dfs中 dfs = [df1, df2, df3] # 替换为你的DataFrame列表 plt.figure(figsize=(10,6)) for idx, df in enumerate(dfs): # 计算相对x轴位置 x = [i/(len(df)-1) for i in range(len(df))] y = df['value'].tolist() plt.plot(x, y, label=f'Series {idx+1}') plt.xlabel('相对位置') plt.ylabel('数值') plt.title('多序列首尾对齐折线图') plt.legend() plt.show()
方案2:时间序列的相对时间归一化
如果你的数据是时间序列,可以将每个序列的时间转换为相对于起始时间的时长占比,同样实现首尾对齐:
示例代码:
import matplotlib.pyplot as plt import pandas as pd dfs = [df1, df2, df3] plt.figure(figsize=(10,6)) for idx, df in enumerate(dfs): # 转换datetime列为时间戳 df['datetime'] = pd.to_datetime(df['datetime']) # 计算每个时间点相对于起始时间的总时长 start_time = df['datetime'].iloc[0] end_time = df['datetime'].iloc[-1] total_duration = (end_time - start_time).total_seconds() # 计算相对时间占比(0到1) x = (df['datetime'] - start_time).dt.total_seconds() / total_duration y = df['value'] plt.plot(x, y, label=f'Time Series {idx+1}') plt.xlabel('相对时间') plt.ylabel('数值') plt.title('时间序列首尾对齐折线图') plt.legend() plt.show()
方案3:数据插值补全(匹配最长序列长度)
如果需要将所有序列调整为相同长度,可以以最长的DataFrame长度为基准,对短序列进行插值补全。这种方法会修改原始数据,但能保持数据的连续性:
示例代码:
import pandas as pd import numpy as np import matplotlib.pyplot as plt dfs = [df1, df2, df3] # 找到最长的序列长度 max_length = max(len(df) for df in dfs) # 对每个序列进行插值补全 adjusted_dfs = [] for df in dfs: # 创建新的索引,长度为max_length new_index = np.linspace(0, len(df)-1, max_length) # 对value列进行插值 adjusted_values = np.interp(new_index, np.arange(len(df)), df['value']) adjusted_dfs.append(pd.Series(adjusted_values)) # 绘制调整后的序列 plt.figure(figsize=(10,6)) for idx, ser in enumerate(adjusted_dfs): plt.plot(ser, label=f'Adjusted Series {idx+1}') plt.xlabel('索引') plt.ylabel('数值') plt.title('插值补全后首尾对齐折线图') plt.legend() plt.show()
内容的提问来源于stack exchange,提问作者FN_
相关产品推荐
相关产品推荐

