Python如何对Pandas DataFrame分组统计并绘制分类5分钟增量变化趋势图
实现方案
核心逻辑说明
- 时间窗口分组:使用pandas的
pd.Grouper按5分钟频率对时间列分组,同时搭配分类字段聚合,得到每个时间窗口下各分类的出现次数 - 历史数据留存:单独维护存储历史统计结果的DataFrame,每次5分钟统计完成后将当前窗口的结果追加进去,用于绘制趋势折线
- 动态绘图:开启matplotlib交互模式,每次统计完成后刷新图表内容,实现每5分钟自动更新
完整可运行代码
import pandas as pd import matplotlib.pyplot as plt from datetime import datetime import time def CreateStats(): print("Reading from file") # 初始化存储所有原始数据的DataFrame df = pd.DataFrame(columns=['time', 'class', 'conf']) # 初始化存储历史统计结果的DataFrame hist_stats = pd.DataFrame(columns=['time_window', 'class', 'count']) pos = 0 # 开启matplotlib交互模式 plt.ion() fig, ax = plt.subplots(figsize=(10, 6)) while True: # 循环执行,每5分钟统计一次 with open("/home/User/Temp/test_data.txt", "r") as fo: for line in fo: splitted = line.split(";") # 替换原right函数实现取最后1位字符,可根据实际需求调整 conf_value = splitted[1].strip()[-1] df.loc[pos] = [datetime.now().strftime("%Y-%m-%d %H:%M:%S"), splitted[0], conf_value] pos += 1 df['time'] = pd.to_datetime(df['time']) # 按5分钟窗口+分类分组统计数量 current_stats = df.groupby([pd.Grouper(key='time', freq='5T'), 'class']).size().reset_index(name='count') current_stats.rename(columns={'time': 'time_window'}, inplace=True) # 追加到历史统计数据中 hist_stats = pd.concat([hist_stats, current_stats], ignore_index=True) # 刷新图表 ax.cla() # 清空上一轮的绘图内容 # 转换为透视表结构:索引为时间窗口,列名为分类,值为对应计数 pivot_df = hist_stats.pivot(index='time_window', columns='class', values='count').fillna(0) # 绘制折线图,每个分类对应一条不同颜色的线 pivot_df.plot(kind='line', marker='o', ax=ax) # 图表样式设置 ax.set_title('各分类5分钟计数变化趋势') ax.set_xlabel('时间窗口') ax.set_ylabel('出现次数') ax.legend(title='分类') plt.xticks(rotation=45) plt.tight_layout() # 渲染更新后的图表 plt.draw() plt.pause(0.1) # 等待5分钟进入下一轮统计 time.sleep(300) if __name__ == "__main__": CreateStats()
关键说明
- 分组统计代码中
freq='5T'代表5分钟时间窗口,可根据需求调整为其他频率,比如10T代表10分钟、1H代表1小时 - 透视表步骤使用
fillna(0)将某时间窗口未出现的分类计数填充为0,避免折线出现断点 - 如果原始文件是增量写入的,可调整逻辑只读取新增内容,避免重复统计历史数据
内容的提问来源于stack exchange,提问作者Magnus_G
相关产品推荐
相关产品推荐

