You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于自定义CropBox提取PDF指定区域文本的问题求助

基于自定义CropBox提取PDF指定区域文本的问题解决

我通过自定义CropBox裁剪PDF后,文件在阅读器中仅显示指定区域,但使用PyPDF2提取文本时仍会获取原页面全部内容,如何实现只提取CropBox指定区域的文本?


现有实现代码

1. 打开PDF并获取第一页

from PyPDF2 import PdfFileReader, PdfFileWriter
from pathlib import Path


pdf_path = (
     Path.home()     
     / "Documents"
     / "XXX"
     / "XXX.pdf" 
            )

pdf = PdfFileReader(str(pdf_path))
numberpages = pdf.getNumPages() # 获取总页数

first_page = pdf.getPage(0)

2. 创建自定义CropBox并设置坐标

# 厘米转英寸:1cm = 0.393700787英寸
inches = 0.393700787

# 厘米单位的坐标(x,y)
lowerLeft_in = (2,8.5)
lowerRight_in = (7,lowerLeft_in[1])
upperLeft_in = (lowerLeft_in[0],9.3)
upperRight_in = (lowerRight_in[0],upperLeft_in[1])

# 转换为PDF的默认单位(1英寸=72点)
lowerLeft = tuple(ti1*72*inches for ti1 in lowerLeft_in)
lowerRight = tuple(ti2*72*inches for ti2 in lowerRight_in)
upperLeft = tuple(ti3*72*inches for ti3 in upperLeft_in)
upperRight = tuple(ti4*72*inches for ti4 in upperRight_in)

first_page.mediaBox.lowerLeft = lowerLeft
first_page.mediaBox.lowerRight = lowerRight
first_page.mediaBox.upperLeft = upperLeft
first_page.mediaBox.upperRight = upperRight

3. 保存裁剪后的PDF

pdf_writer = PdfFileWriter()
pdf_writer.addPage(first_page)
with Path("cropped.pdf").open(mode="wb") as output_file:
    pdf_writer.write(output_file)

4. 文本提取(问题代码)

生成的cropped.pdf在阅读器中显示正常,但执行以下代码仍会提取全页文本:

pdf = PdfFileReader("cropped.pdf")
page = pdf.pages[0]
Newtext = page.extract_text() # 提取PDF全部文本
print(Newtext)

解决方案

PyPDF2的extract_text()方法不会自动识别CropBox过滤文本——裁剪框仅控制显示区域,原PDF的文本内容并未被删除。要提取指定区域的文本,推荐使用pdfplumber库,它支持按坐标区域精准提取文本,步骤如下:

1. 安装pdfplumber

pip install pdfplumber

2. 按CropBox区域提取文本

复用之前定义的坐标,转换为pdfplumber兼容的区域格式(左、上、右、下),注意pdfplumber的坐标原点在页面左上角,需做对应转换:

import pdfplumber

# 复用厘米转英寸规则和原坐标定义
inches = 0.393700787
lowerLeft_in = (2,8.5)
lowerRight_in = (7,lowerLeft_in[1])
upperLeft_in = (lowerLeft_in[0],9.3)

# 转换为pdfplumber的区域参数:(x0, top, x1, bottom),单位为点(1英寸=72点)
# 假设页面为A4纸(高度29.7cm),将原左下角y坐标转换为从顶部开始的距离
x0 = lowerLeft_in[0] * inches * 72
top = (29.7 - upperLeft_in[1]) * inches * 72
x1 = lowerRight_in[0] * inches * 72
bottom = (29.7 - lowerLeft_in[1]) * inches * 72

with pdfplumber.open("cropped.pdf") as pdf:
    page = pdf.pages[0]
    # 提取指定区域内的文本
    cropped_text = page.within_bbox((x0, top, x1, bottom)).extract_text()
    print(cropped_text)

补充说明

  • 若使用非A4尺寸的PDF,需将代码中29.7替换为对应页面的高度(单位:厘米)。
  • 若不想依赖第三方库,也可通过PyPDF2的布局模式提取带位置的文本块,再手动过滤坐标在CropBox内的内容,但实现逻辑复杂,不如pdfplumber高效。

内容的提问来源于stack exchange,提问作者RVF

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 22:57:20