基于DataFrame分组求和生成数据表及可视化图表的技术问询
大型DataFrame的分组统计与可视化实现
我们有一个名为data_frame的大型DataFrame,包含PRE、STATUS、CHR三列,数据示例如下:
PRE STATUS CHR 1_752566 GAINED 1 1_776546 LOST 1 1_832918 NA 1 1_842013 LOST 1 1_846864 GAINED 1 11_8122943 NA 11 11_8188699 GAINED 11 11_8321128 NA 11 23_95137734 NA 23 23_95146814 GAINED 23
1. 按CHR分组统计总行数(生成TOTAL数据表)
使用pandas按CHR分组后统计每组行数,结果保存为TOTAL:
import pandas as pd # 分组统计行数并重置索引,命名统计列为TOTAL_COUNT TOTAL = data_frame.groupby('CHR').size().reset_index(name='TOTAL_COUNT')
输出示例:
CHR TOTAL_COUNT 0 1 5 1 11 3 2 23 2
2. 按CHR分组统计GAINED/LOST行数(生成BY_STATUS数据表)
过滤掉STATUS为NA的记录后,按CHR和STATUS分组统计,再转换为宽表格式:
# 仅保留STATUS为GAINED或LOST的记录,分组统计后转宽表 BY_STATUS = data_frame[data_frame['STATUS'].isin(['GAINED', 'LOST'])] \ .groupby(['CHR', 'STATUS']).size() \ .unstack(fill_value=0) \ .reset_index() # 确保GAINED和LOST列都存在(避免部分分组无对应状态的情况) for col in ['GAINED', 'LOST']: if col not in BY_STATUS.columns: BY_STATUS[col] = 0
输出示例:
CHR GAINED LOST 0 1 2 2 1 11 1 0 2 23 1 0
3. 生成可视化图表
使用matplotlib实现两个统计图表:
3.1 基于TOTAL的总数统计柱状图
import matplotlib.pyplot as plt plt.figure(figsize=(8, 5)) # 绘制柱状图,CHR转为字符串避免数值排序问题 plt.bar(TOTAL['CHR'].astype(str), TOTAL['TOTAL_COUNT'], color='#1f77b4') plt.title('各CHR分组总行数统计') plt.xlabel('CHR编号') plt.ylabel('总行数') plt.grid(axis='y', linestyle='--', alpha=0.7) plt.show()
3.2 基于BY_STATUS的并列对比柱状图
plt.figure(figsize=(8, 5)) bar_width = 0.35 x = range(len(BY_STATUS)) # 绘制并列柱状图 plt.bar([i - bar_width/2 for i in x], BY_STATUS['GAINED'], width=bar_width, label='GAINED', color='#2ca02c') plt.bar([i + bar_width/2 for i in x], BY_STATUS['LOST'], width=bar_width, label='LOST', color='#ff7f0e') plt.title('各CHR分组GAINED与LOST数量对比') plt.xlabel('CHR编号') plt.ylabel('数量') plt.xticks(x, BY_STATUS['CHR'].astype(str)) plt.legend() plt.grid(axis='y', linestyle='--', alpha=0.7) plt.show()
内容的提问来源于stack exchange,提问作者user11924976
相关产品推荐
相关产品推荐

