基于自定义CropBox提取PDF指定区域文本的问题求助
基于自定义CropBox提取PDF指定区域文本的问题解决
我通过自定义CropBox裁剪PDF后,文件在阅读器中仅显示指定区域,但使用PyPDF2提取文本时仍会获取原页面全部内容,如何实现只提取CropBox指定区域的文本?
现有实现代码
1. 打开PDF并获取第一页
from PyPDF2 import PdfFileReader, PdfFileWriter from pathlib import Path pdf_path = ( Path.home() / "Documents" / "XXX" / "XXX.pdf" ) pdf = PdfFileReader(str(pdf_path)) numberpages = pdf.getNumPages() # 获取总页数 first_page = pdf.getPage(0)
2. 创建自定义CropBox并设置坐标
# 厘米转英寸:1cm = 0.393700787英寸 inches = 0.393700787 # 厘米单位的坐标(x,y) lowerLeft_in = (2,8.5) lowerRight_in = (7,lowerLeft_in[1]) upperLeft_in = (lowerLeft_in[0],9.3) upperRight_in = (lowerRight_in[0],upperLeft_in[1]) # 转换为PDF的默认单位(1英寸=72点) lowerLeft = tuple(ti1*72*inches for ti1 in lowerLeft_in) lowerRight = tuple(ti2*72*inches for ti2 in lowerRight_in) upperLeft = tuple(ti3*72*inches for ti3 in upperLeft_in) upperRight = tuple(ti4*72*inches for ti4 in upperRight_in) first_page.mediaBox.lowerLeft = lowerLeft first_page.mediaBox.lowerRight = lowerRight first_page.mediaBox.upperLeft = upperLeft first_page.mediaBox.upperRight = upperRight
3. 保存裁剪后的PDF
pdf_writer = PdfFileWriter() pdf_writer.addPage(first_page) with Path("cropped.pdf").open(mode="wb") as output_file: pdf_writer.write(output_file)
4. 文本提取(问题代码)
生成的cropped.pdf在阅读器中显示正常,但执行以下代码仍会提取全页文本:
pdf = PdfFileReader("cropped.pdf") page = pdf.pages[0] Newtext = page.extract_text() # 提取PDF全部文本 print(Newtext)
解决方案
PyPDF2的extract_text()方法不会自动识别CropBox过滤文本——裁剪框仅控制显示区域,原PDF的文本内容并未被删除。要提取指定区域的文本,推荐使用pdfplumber库,它支持按坐标区域精准提取文本,步骤如下:
1. 安装pdfplumber
pip install pdfplumber
2. 按CropBox区域提取文本
复用之前定义的坐标,转换为pdfplumber兼容的区域格式(左、上、右、下),注意pdfplumber的坐标原点在页面左上角,需做对应转换:
import pdfplumber # 复用厘米转英寸规则和原坐标定义 inches = 0.393700787 lowerLeft_in = (2,8.5) lowerRight_in = (7,lowerLeft_in[1]) upperLeft_in = (lowerLeft_in[0],9.3) # 转换为pdfplumber的区域参数:(x0, top, x1, bottom),单位为点(1英寸=72点) # 假设页面为A4纸(高度29.7cm),将原左下角y坐标转换为从顶部开始的距离 x0 = lowerLeft_in[0] * inches * 72 top = (29.7 - upperLeft_in[1]) * inches * 72 x1 = lowerRight_in[0] * inches * 72 bottom = (29.7 - lowerLeft_in[1]) * inches * 72 with pdfplumber.open("cropped.pdf") as pdf: page = pdf.pages[0] # 提取指定区域内的文本 cropped_text = page.within_bbox((x0, top, x1, bottom)).extract_text() print(cropped_text)
补充说明
- 若使用非A4尺寸的PDF,需将代码中
29.7替换为对应页面的高度(单位:厘米)。 - 若不想依赖第三方库,也可通过PyPDF2的布局模式提取带位置的文本块,再手动过滤坐标在CropBox内的内容,但实现逻辑复杂,不如pdfplumber高效。
内容的提问来源于stack exchange,提问作者RVF
相关产品推荐
相关产品推荐

