多类别时间序列数据可视化:产线批次与阶段趋势图绘制方法
问题
我有一条产线的时间序列数据集,包含batch字段(表示批次名称,字符串类型)和phase字段(表示生产阶段,字符串类型),以datetime作为pandas DataFrame的索引。
需要绘制该时间序列图,要求:
- 叠加各阶段数据,用不同颜色区分不同batch
- 将每个过程变量(
temp1、temp2、press1、press2)放在独立坐标轴上 - 趋势图需基于datetime基准线绘制,确保叠加展示效果
数据集示例
| datetime | temp1 | temp2 | press1 | press2 | batch | phase |
|---|---|---|---|---|---|---|
| 2023-02-03 15:45:34 | 34.45 | 23.34 | 13.23 | 45.5 | 'D' | '10-Wait' |
| ... | ... | ... | ... | ... | 'D' | ... |
| 2023-02-03 15:55:34 | 36.55 | 22.14 | 18.23 | 78.5 | 'D' | '20-Initialise' |
数据集生成代码
import numpy as np import pandas as pd import datetime date = pd.date_range(start='1/1/2023', end='10/06/2023', freq=datetime.timedelta(seconds=30)) tags = ['temp1','temp2','press1','press2'] data=np.random.rand(len(date),len(tags)) df=pd.DataFrame(data,columns=tags).set_index(date) batches = ['A','B','C','D','E','F','G'] n=len(batches) period_start = pd.to_datetime('1/1/2023') period_end = pd.to_datetime('10/06/2023') batch_start = (pd.to_timedelta(np.random.rand(n) * ((period_end - period_start).days + 1), unit='D') + period_start) batch_end = (batch_start + pd.to_timedelta(8,unit='H')) df_batches = pd.DataFrame(data=[batch_start,batch_end],columns=[batches],index=['start','end']).T for item in batches: start_time = df_batches['start'][item] end_time = df_batches['end'][item] df.loc[((df.index>=start_time)&(df.index<=end_time)), 'batch'] = item df.dropna(subset=['batch'],inplace=True) df['phase']='' phases = ['10-Wait','20-Initialise','30-Warm','40-Running'] for batch in batches: wait_len = int(len(df[df['batch']==batch].index)*0.2) init_len = int(len(df[df['batch']==batch].index)*0.4) warm_len = int(len(df[df['batch']==batch].index)*0.6) run_len = int(len(df[df['batch']==batch].index)) wait_start = df[df['batch']==batch].index[0] wait_end = df[df['batch']==batch].index[wait_len] init_end = df[df['batch']==batch].index[init_len] warm_end = df[df['batch']==batch].index[warm_len] run_end = df[df['batch']==batch].index[-1] df['phase'].loc[wait_start:wait_end] = phases[0] df['phase'].loc[wait_end:init_end] = phases[1] df['phase'].loc[init_end:warm_end] = phases[2] df['phase'].loc[warm_end:run_end] = phases[3] df.to_csv('stackoverflowqn.csv')
解决方案
使用matplotlib实现多子图布局,每个子图对应一个过程变量,按批次区分颜色叠加绘制时间序列,具体实现如下:
1. 环境准备
确保已安装依赖库:
pip install pandas numpy matplotlib
2. 完整绘图代码
import pandas as pd import matplotlib.pyplot as plt from matplotlib.colors import ListedColormap # 读取数据集(若使用生成的数据集可直接沿用之前的df) df = pd.read_csv('stackoverflowqn.csv', index_col='datetime', parse_dates=True) # 定义变量列表与批次颜色映射 variables = ['temp1', 'temp2', 'press1', 'press2'] batches = df['batch'].unique() cmap = ListedColormap(plt.cm.tab10.colors[:len(batches)]) batch_color_map = {batch: cmap(i) for i, batch in enumerate(batches)} # 创建共享x轴的多子图布局 fig, axes = plt.subplots(nrows=len(variables), ncols=1, figsize=(12, 10), sharex=True) # 遍历变量绘制各批次曲线 for ax, var in zip(axes, variables): for batch in batches: batch_data = df[df['batch'] == batch] ax.plot(batch_data.index, batch_data[var], label=f'批次 {batch}', color=batch_color_map[batch], alpha=0.7) # 子图样式设置 ax.set_title(f'{var} 趋势') ax.set_ylabel(var) ax.grid(True, alpha=0.3) # 可选:添加生产阶段分隔线 phase_change_points = [] for batch in batches: batch_phases = df[df['batch'] == batch]['phase'] change_indices = batch_phases[batch_phases != batch_phases.shift()].index phase_change_points.extend(change_indices) for ax in axes: for point in phase_change_points: ax.axvline(x=point, color='gray', linestyle='--', alpha=0.5) # 图例与x轴设置 axes[-1].legend(bbox_to_anchor=(1.05, 1), loc='upper left') axes[-1].set_xlabel('时间') # 调整布局避免标签重叠 plt.tight_layout() plt.show()
3. 代码说明
- 共享x轴:通过
sharex=True让所有子图使用同一datetime基准,保证叠加展示的一致性。 - 批次颜色区分:利用
ListedColormap为每个批次分配唯一颜色,避免曲线混淆。 - 阶段标记:自动识别各批次的阶段切换点,添加垂直虚线直观划分生产阶段。
- 图例优化:将图例置于图外,避免遮挡曲线内容。
4. 优化建议
- 若批次数量过多,可降低
alpha参数值(如0.5)提升曲线透明度,减少重叠干扰。 - 可针对
phase字段为同批次曲线的不同阶段设置同色系深浅,进一步区分生产阶段。 - 通过
ax.set_xlim()可指定x轴时间范围,放大查看特定时间段的细节。
内容的提问来源于stack exchange,提问作者hamslice
相关产品推荐
相关产品推荐

