You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无需保存并重新加载文件,能否从PDF指定区域提取文本?

无需保存重新加载PDF,直接提取指定区域文本的方法

可以直接从PDF的指定区域提取文本,无需保存并重新加载文件,以下是两种实用实现方式:

1. 使用PyPDF2实现

通过设置PDF页面的cropbox来限定提取区域,之后直接调用页面的extractText()方法即可获取指定区域文本,无需额外保存文件。示例代码如下:

from PyPDF2 import PdfFileReader

def crop(page, top, left, bottom, right):
    # PyPDF2坐标原点为页面左下角,upper_left对应(left, top),lower_right对应(right, bottom)
    page.cropbox.upper_left = (left, top)
    page.cropbox.lower_right = (right, bottom)
    return page

with open(file_path, 'rb') as pdfFileObj:
    pdfReader = PdfFileReader(pdfFileObj)
    tot_pages = pdfReader.numPages
    print(f"总页数:{tot_pages}")
    # 处理前半部分页面
    for page_num in range(int(tot_pages/2)):
        page = pdfReader.pages[page_num]
        # 传入裁剪区域坐标:top, left, bottom, right
        cropped_page = crop(page, 1000, 170, 210, 1280)
        print(cropped_page.extractText())

注意:原示例代码的坐标参数顺序有误,已修正为符合PyPDF2坐标体系的格式,避免提取区域偏移。

2. 使用tabula-py实现(更简便)

tabula-py支持直接通过X、Y坐标指定提取区域,操作更直观高效,无需手动处理页面裁剪逻辑。示例代码如下:

import tabula

# 指定PDF路径和提取区域,area参数格式为[top, left, bottom, right]
extract_result = tabula.read_pdf(
    file_path,
    pages=range(int(tabula.enumerate_pages(file_path)/2)),  # 处理前半部分页面
    area=[1000, 170, 210, 1280],  # 目标提取区域的坐标
    output_format="json"
)

# 输出提取的文本内容
for item in extract_result:
    for row in item['data']:
        print(" ".join([cell['text'] for cell in row]))

内容的提问来源于stack exchange,提问作者aster94

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:31:12