如何在Python中利用BytesIO内存流基于Matplotlib生成多页PDF报告
Hey there! No need to apologize for asking this—we all run into these little gotchas when starting out, and PDFs have tricky file structures that make straightforward byte concatenation impossible.
The problem with your current code is that you can't just append the raw bytes of two PDF files together. PDFs aren't plain text; they have a structured format with headers, cross-reference tables, and a trailer that tells the reader where to find content. When you write tmp.getvalue() followed by tmp2.getvalue(), you're effectively overwriting the valid PDF structure from the first file with the second one, which is why only the heatmap shows up.
Here are two solid solutions to fix this:
Solution 1: Generate a multi-page PDF directly with Matplotlib (simplest approach)
Matplotlib has a built-in PdfPages class designed exactly for creating multi-page PDFs. This lets you skip the intermediate BytesIO steps entirely and build your PDF page-by-page:
from matplotlib.backends.backend_pdf import PdfPages import seaborn as sns import matplotlib.pyplot as plt # Assuming your DataFrame is named 'data' cols = data.columns.tolist() output_path = "pairgrid.pdf" # Create a PDF file that will hold multiple pages with PdfPages(output_path) as pdf: # First page: PairGrid plot g = sns.PairGrid(data[cols]) g.map_diag(sns.histplot) g.map_upper(sns.scatterplot) g.map_lower(sns.kdeplot, cmap="Set2") pdf.savefig() # Saves the current figure as a new page plt.close() # Clean up the figure to free memory # Second page: Correlation heatmap fig, ax = plt.subplots(figsize=(20,15)) sns.heatmap(data.corr(), cmap='coolwarm', annot=True, fmt=".1f", ax=ax) pdf.savefig(fig) # You can also pass the figure object directly plt.close()
Solution 2: Merge existing PDF BytesIO objects with PyPDF2
If you need to keep your separate BytesIO objects (maybe for other parts of your workflow), you can use the PyPDF2 library to properly merge the two single-page PDFs. First install it via pip:
pip install PyPDF2
Then use this code to merge your streams:
from PyPDF2 import PdfMerger import seaborn as sns import matplotlib.pyplot as plt from io import BytesIO # Your existing code to generate the two individual PDF streams cols = data.columns.tolist() g = sns.PairGrid(data[cols]) g.map_diag(sns.histplot) g.map_upper(sns.scatterplot) g.map_lower(sns.kdeplot, cmap="Set2") tmp = BytesIO() plt.savefig(tmp, format='pdf') plt.close() tmp.seek(0) # Reset the stream to the start tmp2 = BytesIO() fig, ax = plt.subplots(figsize=(20,15)) sns.heatmap(data.corr(), cmap='coolwarm', annot=True, fmt=".1f", ax=ax) plt.savefig(tmp2, format='pdf') plt.close() tmp2.seek(0) # Merge the two PDFs into one multi-page file merger = PdfMerger() merger.append(tmp) # Add the first PDF stream merger.append(tmp2) # Add the second PDF stream # Write the merged PDF to disk output_path = "pairgrid.pdf" with open(output_path, 'wb') as out_file: merger.write(out_file) merger.close()
Either of these methods will give you a two-page PDF with your PairGrid on the first page and heatmap on the second. The first method is more efficient and cleaner, but the second is great if you're working with pre-generated PDF streams.
内容的提问来源于stack exchange,提问作者Ali Mehdi

