在Pandas DataFrame中拼接NumPy数组并绘制带间隙的时间序列图
解决样本不连续的时间序列绘图问题
要在缺失样本的位置显示间隙,核心是为每个测量值生成连续的时间索引,让样本断档的区间自然没有数据,绘图时线条就会自动断开。以下是具体实现步骤:
1. 构造时间索引
每个measurements数组包含7个测量值,假设每个样本内的测量是等间隔的,我们可以为每个测量点生成对应的时间值:以样本编号为起始,每个测量点的时间为 样本编号 + i/7(i从0到6),这样每个样本内的时间是连续的,而样本之间的断档(比如2到5)会形成明显的时间间隔。
2. 合并时间与测量值
将每个样本的时间点和对应测量值配对,合并成一个完整的时间序列DataFrame。
3. 绘制带间隙的时间序列
利用matplotlib绘图时,缺失数据的区间会自动断开线条,呈现间隙效果。
完整代码示例
import pandas as pd import numpy as np import matplotlib.pyplot as plt # 原始DataFrame data = { 'sample': [1, 2, 5, 7], 'measurements': [ np.array([0.2, 0.22, 0.3, 0.7, 0.4, 0.35, 0.2]), np.array([0.2, 0.17, 0.6, 0.6, 0.54, 0.32, 0.2]), np.array([0.2, 0.39, 0.40, 0.53, 0.41, 0.3, 0.2]), np.array([0.2, 0.29, 0.46, 0.68, 0.44, 0.35, 0.2]) ] } df = pd.DataFrame(data) # 生成时间点和测量值列表 time_points = [] measure_values = [] for _, row in df.iterrows(): sample_num = row['sample'] meas_array = row['measurements'] # 生成当前样本的7个时间点:[sample_num, sample_num+1)区间内均分7份 times = np.linspace(sample_num, sample_num + 1, len(meas_array), endpoint=False) time_points.extend(times) measure_values.extend(meas_array) # 构建时间序列DataFrame time_series_df = pd.DataFrame({ 'time': time_points, 'value': measure_values }).sort_values('time').reset_index(drop=True) # 绘图 plt.figure(figsize=(10, 5)) plt.plot(time_series_df['time'], time_series_df['value'], marker='o', linestyle='-') plt.xlabel('Time') plt.ylabel('Measurement Value') plt.title('Time Series with Gaps for Missing Samples') plt.grid(True) plt.show()
效果说明
运行代码后,图表中样本2的最后一个时间点约为2.857,样本5的第一个时间点是5,两者之间无数据,线条会自动断开;同理样本5到7之间也会显示明显间隙,完美呈现数据的断档情况。
内容的提问来源于stack exchange,提问作者MAtennis9
相关产品推荐
相关产品推荐

