如何利用Matplotlib的PdfWriter优化大小DataFrame的PDF格式化
解决DataFrame转PDF时大小适配问题
核心问题分析
原代码固定使用figsize=(12,4),当DataFrame行数较多时,表格会被强制压缩到固定高度的画布中,导致字体过小、内容拥挤。要同时适配小型和大型DataFrame,需要动态调整画布尺寸,并在表格超出单页容量时自动分页。
改进后的代码
import matplotlib.pyplot as plt from matplotlib.backends.backend_pdf import PdfPages def df_to_pdf(panda_df, title, pdf_path): # 基础配置:字体大小、每页最大行数 font_size = 10 max_rows_per_page = 30 num_rows = panda_df.shape[0] num_cols = panda_df.shape[1] # 计算每页需要处理的行范围 page_ranges = [(i, min(i + max_rows_per_page, num_rows)) for i in range(0, num_rows, max_rows_per_page)] with PdfPages(pdf_path) as pdf: for start_idx, end_idx in page_ranges: # 动态计算画布高度:每行约占0.4单位高度,加上标题和边距 fig_height = max(4, (end_idx - start_idx) * 0.4 + 2) fig, ax = plt.subplots(figsize=(12, fig_height)) ax.axis('tight') ax.axis('off') ax.set_title(title, fontsize=font_size+2, pad=15) # 截取当前页的DataFrame数据 page_df = panda_df.iloc[start_idx:end_idx] # 创建表格,设置自动列宽、字体大小 table = ax.table( cellText=page_df.values, colLabels=panda_df.columns, loc='center', cellLoc='center', fontsize=font_size ) # 自动调整列宽:根据内容长度设置列宽比例 table.auto_set_column_width(col=list(range(num_cols))) # 优化表格样式 table.scale(1, 1.2) # 调整行高,避免内容拥挤 pdf.savefig(fig, bbox_inches='tight', dpi=150) plt.close(fig)
关键优化点
- 动态画布高度:根据当前页的行数计算
fig_height,小型DataFrame保持最小高度4,大型DataFrame按行数扩展高度,避免压缩。 - 自动分页:设置
max_rows_per_page控制每页最大行数,超过时自动拆分多页。 - 表格样式优化:
- 用
auto_set_column_width自动适配列宽,避免列内容被截断。 - 调整
table.scale()增加行高,提升可读性。 - 统一字体大小,标题略大于表格内容,保持视觉层次。
- 用
- 资源清理:每页生成后关闭画布,避免内存占用过高。
使用示例
# 假设你的DataFrame是df,标题为"业务数据统计" df_to_pdf(df, "业务数据统计", "output.pdf")
内容的提问来源于stack exchange,提问作者lunbox
相关产品推荐
相关产品推荐

