You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Seaborn/Pandas图表识别引发CPU/IO峰值的批处理作业

分析引发CPU/IO峰值的作业类型的简便可视化方案

核心思路

不用逐个计算单个作业与CPU/IO的相关性,而是通过聚焦峰值时刻的作业分布,直接关联两者的对应关系,大幅简化分析流程。

步骤1:先标记峰值数据

先筛选出CPU或IO处于峰值状态的记录(可按百分比阈值或固定阈值筛选):

import pandas as pd

# 假设你的原始数据已转为DataFrame
data = [['Timestamp', 'CPU%', 'IO', 'Job1', 'Job2', 'Job3'], 
        ['2022-08-06 10:31:59.233', '10', '90', 1, 0, 0], 
        ['2022-08-06 10:32:19.235', '30', '40', 1, 4, 2]]
df = pd.DataFrame(data[1:], columns=data[0])
# 转换数值类型
df[['CPU%', 'IO']] = df[['CPU%', 'IO']].astype(int)

# 取前5%的高值作为峰值阈值(可根据业务调整)
cpu_threshold = df['CPU%'].quantile(0.95)
io_threshold = df['IO'].quantile(0.95)

# 筛选出CPU或IO达峰的时刻数据
peak_data = df[(df['CPU%'] >= cpu_threshold) | (df['IO'] >= io_threshold)]

步骤2:可视化峰值时刻的作业特征

方法1:堆叠柱状图(Pandas)

直观展示峰值时刻各作业的运行数量分布,一眼定位高频/高数量的作业:

import matplotlib.pyplot as plt

# 提取所有作业列
job_cols = [col for col in df.columns if col.startswith('Job')]

# 绘制堆叠柱状图
peak_data[job_cols].plot(kind='bar', stacked=True, figsize=(12,6))
plt.title('CPU/IO峰值时刻的作业运行数量分布')
plt.xlabel('时间戳')
plt.ylabel('作业运行数量')
plt.legend(bbox_to_anchor=(1.05, 1), loc='upper left')
plt.tight_layout()
plt.show()

方法2:分组小提琴图(Seaborn)

对比峰值时刻与非峰值时刻各作业的运行数量分布,清晰看出哪些作业在峰值时活跃度显著提升:

import seaborn as sns

# 将宽表转为长表,适配Seaborn绘图格式
df_melt = df.melt(id_vars=['Timestamp', 'CPU%', 'IO'], 
                  value_vars=job_cols, 
                  var_name='作业类型', 
                  value_name='运行数量')

# 添加"是否峰值"标记列
df_melt['是否峰值'] = df_melt.apply(
    lambda x: '是' if (x['CPU%'] >= cpu_threshold) | (x['IO'] >= io_threshold) else '否', 
    axis=1
)

# 绘制小提琴图
plt.figure(figsize=(15,8))
sns.violinplot(data=df_melt, x='作业类型', y='运行数量', hue='是否峰值', split=True)
plt.xticks(rotation=90)
plt.title('各作业在峰值/非峰值时刻的运行数量分布')
plt.tight_layout()
plt.show()

方法3:相关性热力图(Seaborn)

仅聚焦作业列与CPU/IO的相关性,避免全量计算的繁琐:

# 计算作业列与CPU、IO的相关性矩阵
corr_matrix = df[job_cols + ['CPU%', 'IO']].corr()
# 只保留作业与CPU/IO的对应关系
corr_target = corr_matrix.loc[job_cols, ['CPU%', 'IO']]

# 绘制热力图
plt.figure(figsize=(10,12))
sns.heatmap(corr_target, annot=True, cmap='coolwarm', fmt='.2f')
plt.title('作业类型与CPU/IO的相关性热力图')
plt.tight_layout()
plt.show()

简化分析的小技巧

  • 先过滤掉全时段运行数量为0的作业,减少绘图复杂度
  • 用peak_data[job_cols].sum().sort_values(ascending=False)快速统计峰值时刻各作业的总运行量,先定位重点嫌疑作业再做可视化

内容的提问来源于stack exchange,提问作者kamal kant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 06:36:35