使用Python3 Matplotlib批量绘制多CSV文件分组柱状图需求
解决方案:CSV文件柱状图绘制(单文件+批量子图)
1. 单文件柱状图实现
以下代码完整实现单文件绘图,并包含所有样式调整需求:
import matplotlib.pyplot as plt import pandas as pd # 读取目标CSV文件 file_path = "af-dist-barplot/chr10_af_dist.csv" df = pd.read_csv(file_path) # 创建画布 plt.figure(figsize=(12, 6)) # 自定义柱子宽度与间距 bar_width = 0.35 x_indices = range(len(df['bin'])) # 绘制两组柱状图,指定自定义颜色 plt.bar([i - bar_width/2 for i in x_indices], df['1KGP'], width=bar_width, label='1KGP', color='#1f77b4') plt.bar([i + bar_width/2 for i in x_indices], df['10KGP'], width=bar_width, label='10KGP', color='#ff7f0e') # X轴样式调整:设置bin值为刻度、缩小字体、垂直显示 plt.xticks(x_indices, df['bin'], fontsize=8, rotation=90) # 设置轴标签与标题 plt.xlabel('Bin', fontsize=10) plt.ylabel('Count', fontsize=10) plt.title(f'Frequency Distribution - {file_path.split("/")[-1]}', fontsize=12) plt.legend() # 自动调整布局,避免标签截断 plt.tight_layout() # 保存并显示图片 plt.savefig(f'{file_path.split("/")[-1].replace(".csv", ".png")}', dpi=300) plt.show() plt.close()
样式调整说明
- X轴设置:通过
fontsize调整刻度字体大小,rotation=90实现垂直显示 - 柱宽与间距:用
bar_width定义柱子宽度,通过偏移X坐标(i ± bar_width/2)控制两组柱子的间距 - 自定义颜色:在
plt.bar()中通过color参数指定,支持十六进制色码或颜色名称
2. 批量文件子图整合实现
以下代码将所有CSV文件的图表整合到一个画布中,每个子图以对应文件名作为标题:
from pathlib import Path import matplotlib.pyplot as plt import pandas as pd # 目标文件夹路径 folder_path = Path('af_dist-barplot') csv_files = list(folder_path.glob('*.csv')) # 设置子图布局:每行4个,自动计算所需行数 cols = 4 rows = (len(csv_files) + cols - 1) // cols # 向上取整计算行数 # 创建画布与子图 fig, axes = plt.subplots(rows, cols, figsize=(cols*6, rows*4)) axes = axes.flatten() # 将二维轴数组转为一维,方便遍历 # 统一样式参数 bar_width = 0.35 color_1kgp = '#1f77b4' color_10kgp = '#ff7f0e' # 遍历每个文件绘制子图 for idx, file in enumerate(csv_files): ax = axes[idx] df = pd.read_csv(file) x_indices = range(len(df['bin'])) # 绘制两组柱状图 ax.bar([i - bar_width/2 for i in x_indices], df['1KGP'], width=bar_width, color=color_1kgp) ax.bar([i + bar_width/2 for i in x_indices], df['10KGP'], width=bar_width, color=color_10kgp) # 设置子图标题为文件名 ax.set_title(file.name, fontsize=10) # X轴样式调整 ax.set_xticks(x_indices) ax.set_xticklabels(df['bin'], fontsize=6, rotation=90) # 设置轴标签 ax.set_xlabel('Bin', fontsize=8) ax.set_ylabel('Count', fontsize=8) # 仅在第一个子图添加图例,避免重复 if idx == 0: ax.legend(['1KGP', '10KGP'], fontsize=8) # 隐藏多余的空轴(当文件数不是列数整数倍时) for idx in range(len(csv_files), len(axes)): axes[idx].axis('off') # 调整子图间距,避免内容重叠 plt.tight_layout() # 保存整图 plt.savefig('all_af_dist_plots.png', dpi=300) plt.show() plt.close()
关键细节说明
- 子图布局:自动计算所需行数,确保所有文件都能分配到子图
- 标题设置:直接使用
file.name获取文件名作为子图标题 - 冗余处理:隐藏未使用的空轴,优化整体布局
- 图例优化:仅在第一个子图显示图例,减少视觉冗余
可选优化建议
- 如果bin值过多导致X轴拥挤,可通过
ax.set_xticks(x_indices[::2])只显示间隔刻度 - 若两组数据数值差异极大,可添加
ax.set_yscale('log')将Y轴设为对数刻度,提升可读性
内容的提问来源于stack exchange,提问作者user3214212
相关产品推荐
相关产品推荐

