如何用Python彻底裁剪PDF内容而非仅设置CropBox?
如何彻底裁剪PDF区域并删除Bounding Box外的内容?
我尝试编写脚本裁剪PDF的指定区域,合并为单页新PDF,但修改CropBox后,被裁剪的数据仅被隐藏而非删除,导致后续用文本解析器处理输出PDF时,仍能提取到隐藏文本(无需OCR)。
请问如何裁剪页面并彻底删除bounding box外的数据?
示例说明
在目标PDF中,我想要裁剪两个蓝色框区域并合并为单页输出文件,但操作后选中文本时仍包含隐藏内容,PDF解析器也会读取到隐藏文本。



我最初以为是PyMuPDF库的问题,但使用PyPDF2编写等效代码后仍遇到相同问题。以下是我尝试的代码:
尝试过的代码
PyMuPDF代码
from fitz import Document, Page, Rect # 定义要提取的区域列表,每个区域包含页码和矩形坐标 boxes = [ { 'page_number': 0, 'rect': Rect(0, 54, 595, 189) }, { 'page_number': 0, 'rect': Rect(0, 317, 595, 459) } ] # 计算新页面尺寸:最大宽度 + 所有区域高度之和 new_page_rect = Rect( 0, 0, max([box['rect'].width for box in boxes]) + 1, sum([box['rect'].height for box in boxes]) + 1 ) # 打开输入PDF并创建输出PDF with Document(r"lorem_ipsum.pdf") as input_document, Document() as output_document: # 创建新页面 new_page: Page = output_document.new_page( width=new_page_rect.width, height=new_page_rect.height ) last_y_coord = 0 for box in boxes: # 复制原页面到输入文档末尾 input_document.copy_page(box['page_number']) page = input_document[-1] # 设置裁剪框 page.set_cropbox(box['rect']) # 计算当前区域在新页面的位置 rect = Rect( 0, last_y_coord, box['rect'].width, last_y_coord + box['rect'].height, ) last_y_coord = rect.y1 + 1 # 将裁剪后的页面绘制到新页面 new_page.show_pdf_page(rect, input_document, page.number) # 保存输出PDF output_document.save(filename=r"output_PyMuPDF.pdf", garbage=3, deflate=True, pretty=True)
PyPDF2代码
import io import PyPDF2 from PyPDF2 import Transformation from copy import copy # 定义要提取的区域列表 boxes = [ { 'page_number': 0, 'rect': (0, 54, 595, 189) }, { 'page_number': 0, 'rect': (0, 317, 595, 459) } ] # 计算新页面尺寸 new_page_width = max([box['rect'][2] - box['rect'][0] for box in boxes]) + 1 new_page_height = sum([box['rect'][3] - box['rect'][1] for box in boxes]) + 1 # 打开输入和输出文件 with open(r"lorem_ipsum.pdf", "rb") as input_file, open(r"output_PyPDF2.pdf", "wb") as output_file: reader = PyPDF2.PdfFileReader(input_file) writer = PyPDF2.PdfFileWriter() # 创建空白新页面 new_page = PyPDF2.PageObject.create_blank_page( pdf=None, width=new_page_width, height=new_page_height ) last_y_coord = new_page_height for box in boxes: # 复制原页面 page = copy(reader.getPage(box['page_number'])) page_height = page.mediabox.upper_right[1] # 转换坐标(PyPDF2使用左下角为原点) x0 = box['rect'][0] y0 = page_height - box['rect'][3] x1 = box['rect'][2] y1 = page_height - box['rect'][1] # 计算平移变换 tx = -x0 ty = last_y_coord - y1 transformation = Transformation().translate(tx=tx, ty=ty) page.add_transformation(transformation) # 更新裁剪框 page.cropbox.lower_left = (x0, y0 + ty) page.cropbox.upper_right = (x1, y1 + ty) # 合并到新页面 new_page.merge_page(page) last_y_coord -= (y1 - y0 + 1) # 添加新页面并保存 writer.addPage(new_page) writer.write(output_file)

解决方案:彻底移除框外内容
单纯修改CropBox/MediaBox只是修改PDF的显示区域,不会删除页面内的实际内容,所以文本解析器仍能读取到隐藏内容。要彻底删除,需要提取指定区域内的内容并重新构建页面,或使用底层工具进行物理裁剪。
方法1:使用PyMuPDF提取区域内容并重建页面
直接提取指定矩形内的文本、图像等元素,在新页面上重新添加,确保只有框内内容被保留:
from fitz import Document, Rect boxes = [ {'page_number': 0, 'rect': Rect(0, 54, 595, 189)}, {'page_number': 0, 'rect': Rect(0, 317, 595, 459)} ] # 计算新页面尺寸 new_width = max(box['rect'].width for box in boxes) + 1 new_height = sum(box['rect'].height for box in boxes) + 1 with Document("lorem_ipsum.pdf") as input_doc, Document() as output_doc: new_page = output_doc.new_page(width=new_width, height=new_height) current_y = 0 for box in boxes: src_page = input_doc[box['page_number']] crop_rect = box['rect'] # 提取矩形内的文本块 text_blocks = src_page.get_text("blocks", clip=crop_rect) # 提取矩形内的图像 images = src_page.get_images(clip=crop_rect) # 添加文本到新页面 for block in text_blocks: # 原文本块坐标 block_rect = Rect(block[0], block[1], block[2], block[3]) # 转换到新页面的坐标 new_block_rect = Rect( block_rect.x0 - crop_rect.x0, current_y + (block_rect.y0 - crop_rect.y0), block_rect.x1 - crop_rect.x0, current_y + (block_rect.y1 - crop_rect.y0) ) # 插入文本(可根据实际情况调整字体大小等参数) new_page.insert_textbox( new_block_rect, block[4], fontsize=12, align=0 ) # 添加图像到新页面 for img_info in images: img = src_page.get_image(img_info[0]) img_rect = src_page.get_image_bbox(img_info) # 转换到新页面的坐标 new_img_rect = Rect( img_rect.x0 - crop_rect.x0, current_y + (img_rect.y0 - crop_rect.y0), img_rect.x1 - crop_rect.x0, current_y + (img_rect.y1 - crop_rect.y0) ) # 插入图像 new_page.insert_image(new_img_rect, stream=img['image']) # 更新当前y坐标,预留间距 current_y += crop_rect.height + 1 # 保存清理后的PDF output_doc.save("output_clean.pdf", garbage=3, deflate=True)
方法2:使用Ghostscript进行物理裁剪(命令行)
通过Ghostscript直接对PDF进行底层裁剪,彻底删除框外内容:
# 裁剪第一个区域到临时文件 gs -o temp1.pdf -sDEVICE=pdfwrite -c "[/CropBox [0 54 595 189] /PAGES pdfmark" -f lorem_ipsum.pdf # 裁剪第二个区域到临时文件 gs -o temp2.pdf -sDEVICE=pdfwrite -c "[/CropBox [0 317 595 459] /PAGES pdfmark" -f lorem_ipsum.pdf # 合并两个临时文件为单页 gs -o final_output.pdf -sDEVICE=pdfwrite -dNOPAUSE -dBATCH -dPDFFitPage temp1.pdf temp2.pdf
补充说明
- 如果不需要保留文本可编辑性,也可以将指定区域转为图片后插入新页面(使用
src_page.get_pixmap(clip=crop_rect)获取图像,再new_page.insert_image),这种方式同样能彻底清除隐藏文本,但文本会变为图像格式。
内容的提问来源于stack exchange,提问作者Igor Micadei
相关产品推荐
相关产品推荐

