You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用PyPDF2与PyMuPDF(fitz)将A4 PDF分割为A6四等分并过滤空白页面

Fixing Blank Page Issue When Splitting A4 PDF to A6 with PyMuPDF

Hey there! Let's work through this problem you're having with splitting A4 PDFs into A6-sized pages without saving those empty ones. I’ve looked over your code, and the core issue is that right now you’re generating a page for every quadrant—even if there’s no visible content in it. Plus, your post-split detection isn’t working because blank pages might still carry underlying PDF metadata or empty stream objects that trick get_text() into thinking there’s content, even though nothing shows up in Acrobat.

The Solution: Check for Content Before Creating Pages

Instead of generating all four pages and trying to delete blanks later, we can check each A6-sized quadrant in the original A4 page before creating a new page. PyMuPDF has built-in tools to verify if a specific region has text or graphical content—we’ll use those to skip empty regions entirely.

Here’s the revised fitz_four_piaces function with blank detection built in:

def fitz_four_piaces(input_file, output_file):
    input_file = str(input_file.absolute())
    src = fitz.open(input_file)
    doc = fitz.open()  # Empty output PDF

    for spage in src:  # Iterate over each page in input
        r = spage.rect  # Full page rectangle
        d = fitz.Rect(spage.cropbox_position, spage.cropbox_position)  # CropBox offset

        # Define the four A6 quadrants
        r1 = r / 2  # Top-left
        r2 = r1 + (r1.width, 0, r1.width, 0)  # Top-right
        r3 = r1 + (0, r1.height, 0, r1.height)  # Bottom-left
        r4 = fitz.Rect(r1.br, r.br)  # Bottom-right
        rect_list = [r1, r2, r3, r4]

        for rx in rect_list:
            rx += d  # Apply CropBox offset to the region
            # Check if this region has any content
            has_text = bool(spage.get_textbox(rx).strip())
            has_drawings = bool(spage.get_drawings(clip=rx))
            
            # Only create a page if there's text or drawings
            if has_text or has_drawings:
                page = doc.new_page(-1, width=rx.width, height=rx.height)
                page.show_pdf_page(
                    page.rect,
                    src,
                    spage.number,
                    clip=rx
                )
    
    # Save the output with compression
    doc.save(output_file, deflate=True)

Key Changes Explained

  1. Content Detection:
    • spage.get_textbox(rx).strip(): Pulls all text within the quadrant, strips whitespace, and checks if anything meaningful remains.
    • spage.get_drawings(clip=rx): Looks for graphical elements (lines, shapes, images) in the quadrant. A non-empty list means the region has visible content.
  2. Conditional Page Creation: We only generate a new page if either text or drawings exist in the quadrant—skipping empty regions entirely.

Why Your Previous Detection Failed

When you split pages using show_pdf_page, even empty quadrants create pages that inherit some of the original PDF’s structure (like empty stream objects). These don’t show up visually, but get_text() might pick up empty strings or metadata, making it seem like there’s content when there isn’t. Checking the original page’s region directly avoids this false positive issue.

Give this revised function a try—it should only save A6 pages that actually have content, and you won’t have to deal with those ghost blank pages anymore!

内容的提问来源于stack exchange,提问作者Max

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 19:17:27