You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python彻底裁剪PDF内容而非仅设置CropBox?

如何彻底裁剪PDF区域并删除Bounding Box外的内容?

我尝试编写脚本裁剪PDF的指定区域,合并为单页新PDF,但修改CropBox后,被裁剪的数据仅被隐藏而非删除,导致后续用文本解析器处理输出PDF时,仍能提取到隐藏文本(无需OCR)。

请问如何裁剪页面并彻底删除bounding box外的数据?

示例说明

在目标PDF中,我想要裁剪两个蓝色框区域并合并为单页输出文件,但操作后选中文本时仍包含隐藏内容,PDF解析器也会读取到隐藏文本。

要裁剪合并的Bounding Boxes
裁剪后的单页输出文件选中文本情况
PDF解析器读取到隐藏文本

我最初以为是PyMuPDF库的问题,但使用PyPDF2编写等效代码后仍遇到相同问题。以下是我尝试的代码:


尝试过的代码

PyMuPDF代码

from fitz import Document, Page, Rect

# 定义要提取的区域列表,每个区域包含页码和矩形坐标
boxes = [
    {
        'page_number': 0,
        'rect': Rect(0, 54, 595, 189)
    },
    {
        'page_number': 0,
        'rect': Rect(0, 317, 595, 459)
    }
]

# 计算新页面尺寸:最大宽度 + 所有区域高度之和
new_page_rect = Rect(
    0,
    0,
    max([box['rect'].width for box in boxes]) + 1,
    sum([box['rect'].height for box in boxes]) + 1
)

# 打开输入PDF并创建输出PDF
with Document(r"lorem_ipsum.pdf") as input_document, Document() as output_document:

    # 创建新页面
    new_page: Page = output_document.new_page(
        width=new_page_rect.width,
        height=new_page_rect.height
    )

    last_y_coord = 0

    for box in boxes:
        # 复制原页面到输入文档末尾
        input_document.copy_page(box['page_number'])
        page = input_document[-1]
        # 设置裁剪框
        page.set_cropbox(box['rect'])

        # 计算当前区域在新页面的位置
        rect = Rect(
            0,
            last_y_coord,
            box['rect'].width,
            last_y_coord + box['rect'].height,
        )

        last_y_coord = rect.y1 + 1

        # 将裁剪后的页面绘制到新页面
        new_page.show_pdf_page(rect, input_document, page.number)

    # 保存输出PDF
    output_document.save(filename=r"output_PyMuPDF.pdf", garbage=3, deflate=True, pretty=True)

PyPDF2代码

import io
import PyPDF2
from PyPDF2 import Transformation
from copy import copy

# 定义要提取的区域列表
boxes = [
    {
        'page_number': 0,
        'rect': (0, 54, 595, 189)
    },
    {
        'page_number': 0,
        'rect': (0, 317, 595, 459)
    }
]

# 计算新页面尺寸
new_page_width = max([box['rect'][2] - box['rect'][0] for box in boxes]) + 1
new_page_height = sum([box['rect'][3] - box['rect'][1] for box in boxes]) + 1

# 打开输入和输出文件
with open(r"lorem_ipsum.pdf", "rb") as input_file, open(r"output_PyPDF2.pdf", "wb") as output_file:

    reader = PyPDF2.PdfFileReader(input_file)
    writer = PyPDF2.PdfFileWriter()

    # 创建空白新页面
    new_page = PyPDF2.PageObject.create_blank_page(
        pdf=None,
        width=new_page_width,
        height=new_page_height
    )

    last_y_coord = new_page_height

    for box in boxes:
        # 复制原页面
        page = copy(reader.getPage(box['page_number']))
        page_height = page.mediabox.upper_right[1]

        # 转换坐标(PyPDF2使用左下角为原点)
        x0 = box['rect'][0]
        y0 = page_height - box['rect'][3]
        x1 = box['rect'][2]
        y1 = page_height - box['rect'][1]

        # 计算平移变换
        tx = -x0
        ty = last_y_coord - y1

        transformation = Transformation().translate(tx=tx, ty=ty)
        page.add_transformation(transformation)

        # 更新裁剪框
        page.cropbox.lower_left = (x0, y0 + ty)
        page.cropbox.upper_right = (x1, y1 + ty)

        # 合并到新页面
        new_page.merge_page(page)

        last_y_coord -= (y1 - y0 + 1)

    # 添加新页面并保存
    writer.addPage(new_page)
    writer.write(output_file)

使用PyPDF2裁剪后的单页输出文件选中文本情况


解决方案:彻底移除框外内容

单纯修改CropBox/MediaBox只是修改PDF的显示区域,不会删除页面内的实际内容,所以文本解析器仍能读取到隐藏内容。要彻底删除,需要提取指定区域内的内容并重新构建页面,或使用底层工具进行物理裁剪。

方法1:使用PyMuPDF提取区域内容并重建页面

直接提取指定矩形内的文本、图像等元素,在新页面上重新添加,确保只有框内内容被保留:

from fitz import Document, Rect

boxes = [
    {'page_number': 0, 'rect': Rect(0, 54, 595, 189)},
    {'page_number': 0, 'rect': Rect(0, 317, 595, 459)}
]

# 计算新页面尺寸
new_width = max(box['rect'].width for box in boxes) + 1
new_height = sum(box['rect'].height for box in boxes) + 1

with Document("lorem_ipsum.pdf") as input_doc, Document() as output_doc:
    new_page = output_doc.new_page(width=new_width, height=new_height)
    current_y = 0

    for box in boxes:
        src_page = input_doc[box['page_number']]
        crop_rect = box['rect']
        
        # 提取矩形内的文本块
        text_blocks = src_page.get_text("blocks", clip=crop_rect)
        # 提取矩形内的图像
        images = src_page.get_images(clip=crop_rect)

        # 添加文本到新页面
        for block in text_blocks:
            # 原文本块坐标
            block_rect = Rect(block[0], block[1], block[2], block[3])
            # 转换到新页面的坐标
            new_block_rect = Rect(
                block_rect.x0 - crop_rect.x0,
                current_y + (block_rect.y0 - crop_rect.y0),
                block_rect.x1 - crop_rect.x0,
                current_y + (block_rect.y1 - crop_rect.y0)
            )
            # 插入文本(可根据实际情况调整字体大小等参数)
            new_page.insert_textbox(
                new_block_rect,
                block[4],
                fontsize=12,
                align=0
            )

        # 添加图像到新页面
        for img_info in images:
            img = src_page.get_image(img_info[0])
            img_rect = src_page.get_image_bbox(img_info)
            # 转换到新页面的坐标
            new_img_rect = Rect(
                img_rect.x0 - crop_rect.x0,
                current_y + (img_rect.y0 - crop_rect.y0),
                img_rect.x1 - crop_rect.x0,
                current_y + (img_rect.y1 - crop_rect.y0)
            )
            # 插入图像
            new_page.insert_image(new_img_rect, stream=img['image'])

        # 更新当前y坐标,预留间距
        current_y += crop_rect.height + 1

    # 保存清理后的PDF
    output_doc.save("output_clean.pdf", garbage=3, deflate=True)

方法2:使用Ghostscript进行物理裁剪(命令行)

通过Ghostscript直接对PDF进行底层裁剪,彻底删除框外内容:

# 裁剪第一个区域到临时文件
gs -o temp1.pdf -sDEVICE=pdfwrite -c "[/CropBox [0 54 595 189] /PAGES pdfmark" -f lorem_ipsum.pdf
# 裁剪第二个区域到临时文件
gs -o temp2.pdf -sDEVICE=pdfwrite -c "[/CropBox [0 317 595 459] /PAGES pdfmark" -f lorem_ipsum.pdf
# 合并两个临时文件为单页
gs -o final_output.pdf -sDEVICE=pdfwrite -dNOPAUSE -dBATCH -dPDFFitPage temp1.pdf temp2.pdf

补充说明

  • 如果不需要保留文本可编辑性,也可以将指定区域转为图片后插入新页面(使用src_page.get_pixmap(clip=crop_rect)获取图像,再new_page.insert_image),这种方式同样能彻底清除隐藏文本,但文本会变为图像格式。

内容的提问来源于stack exchange,提问作者Igor Micadei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 04:35:56