如何使用PyPDF2与PyMuPDF(fitz)将A4 PDF分割为A6四等分并过滤空白页面
Hey there! Let's work through this problem you're having with splitting A4 PDFs into A6-sized pages without saving those empty ones. I’ve looked over your code, and the core issue is that right now you’re generating a page for every quadrant—even if there’s no visible content in it. Plus, your post-split detection isn’t working because blank pages might still carry underlying PDF metadata or empty stream objects that trick get_text() into thinking there’s content, even though nothing shows up in Acrobat.
The Solution: Check for Content Before Creating Pages
Instead of generating all four pages and trying to delete blanks later, we can check each A6-sized quadrant in the original A4 page before creating a new page. PyMuPDF has built-in tools to verify if a specific region has text or graphical content—we’ll use those to skip empty regions entirely.
Here’s the revised fitz_four_piaces function with blank detection built in:
def fitz_four_piaces(input_file, output_file): input_file = str(input_file.absolute()) src = fitz.open(input_file) doc = fitz.open() # Empty output PDF for spage in src: # Iterate over each page in input r = spage.rect # Full page rectangle d = fitz.Rect(spage.cropbox_position, spage.cropbox_position) # CropBox offset # Define the four A6 quadrants r1 = r / 2 # Top-left r2 = r1 + (r1.width, 0, r1.width, 0) # Top-right r3 = r1 + (0, r1.height, 0, r1.height) # Bottom-left r4 = fitz.Rect(r1.br, r.br) # Bottom-right rect_list = [r1, r2, r3, r4] for rx in rect_list: rx += d # Apply CropBox offset to the region # Check if this region has any content has_text = bool(spage.get_textbox(rx).strip()) has_drawings = bool(spage.get_drawings(clip=rx)) # Only create a page if there's text or drawings if has_text or has_drawings: page = doc.new_page(-1, width=rx.width, height=rx.height) page.show_pdf_page( page.rect, src, spage.number, clip=rx ) # Save the output with compression doc.save(output_file, deflate=True)
Key Changes Explained
- Content Detection:
spage.get_textbox(rx).strip(): Pulls all text within the quadrant, strips whitespace, and checks if anything meaningful remains.spage.get_drawings(clip=rx): Looks for graphical elements (lines, shapes, images) in the quadrant. A non-empty list means the region has visible content.
- Conditional Page Creation: We only generate a new page if either text or drawings exist in the quadrant—skipping empty regions entirely.
Why Your Previous Detection Failed
When you split pages using show_pdf_page, even empty quadrants create pages that inherit some of the original PDF’s structure (like empty stream objects). These don’t show up visually, but get_text() might pick up empty strings or metadata, making it seem like there’s content when there isn’t. Checking the original page’s region directly avoids this false positive issue.
Give this revised function a try—it should only save A6 pages that actually have content, and you won’t have to deal with those ghost blank pages anymore!
内容的提问来源于stack exchange,提问作者Max

