无需保存并重新加载文件,能否从PDF指定区域提取文本?
无需保存重新加载PDF,直接提取指定区域文本的方法
可以直接从PDF的指定区域提取文本,无需保存并重新加载文件,以下是两种实用实现方式:
1. 使用PyPDF2实现
通过设置PDF页面的cropbox来限定提取区域,之后直接调用页面的extractText()方法即可获取指定区域文本,无需额外保存文件。示例代码如下:
from PyPDF2 import PdfFileReader def crop(page, top, left, bottom, right): # PyPDF2坐标原点为页面左下角,upper_left对应(left, top),lower_right对应(right, bottom) page.cropbox.upper_left = (left, top) page.cropbox.lower_right = (right, bottom) return page with open(file_path, 'rb') as pdfFileObj: pdfReader = PdfFileReader(pdfFileObj) tot_pages = pdfReader.numPages print(f"总页数:{tot_pages}") # 处理前半部分页面 for page_num in range(int(tot_pages/2)): page = pdfReader.pages[page_num] # 传入裁剪区域坐标:top, left, bottom, right cropped_page = crop(page, 1000, 170, 210, 1280) print(cropped_page.extractText())
注意:原示例代码的坐标参数顺序有误,已修正为符合PyPDF2坐标体系的格式,避免提取区域偏移。
2. 使用tabula-py实现(更简便)
tabula-py支持直接通过X、Y坐标指定提取区域,操作更直观高效,无需手动处理页面裁剪逻辑。示例代码如下:
import tabula # 指定PDF路径和提取区域,area参数格式为[top, left, bottom, right] extract_result = tabula.read_pdf( file_path, pages=range(int(tabula.enumerate_pages(file_path)/2)), # 处理前半部分页面 area=[1000, 170, 210, 1280], # 目标提取区域的坐标 output_format="json" ) # 输出提取的文本内容 for item in extract_result: for row in item['data']: print(" ".join([cell['text'] for cell in row]))
内容的提问来源于stack exchange,提问作者aster94
相关产品推荐
相关产品推荐

