You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用PyPDF2自动识别PDF页眉页脚尺寸以提取正文?

基于PyPDF2自动识别页眉页脚区域提取正文

问题背景

我正在使用PyPDF2读取PDF文件并转换为文本格式,但现有代码中固定的y坐标阈值(50和720)仅对部分PDF有效。希望基于PyPDF2实现首次读取时自动提取每份文件的页眉、页脚尺寸,替换固定值来过滤正文。

原测试代码:

from PyPDF2 import PdfReader

# 替换为不同的PDF文件名
reader = PdfReader("GeoBase_NHNC1_Data_Model_UML_EN.pdf")

page = reader.pages[3]

parts = []


def visitor_body(text, cm, tm, fontDict, fontSize):
    y = tm[5]
    if y > 50 and y < 720:
        parts.append(text)


page.extract_text(visitor_text=visitor_body)
text_body = "".join(parts)

print(text_body)

解决方案

实现思路

自动识别页眉页脚的核心是先收集页面文本的y坐标分布,通过统计找出重复出现的顶部(页眉)和底部(页脚)区域边界,再用动态计算的边界过滤正文:

  1. 选取PDF的若干样本页面(如前3页),收集所有非空文本块的y坐标
  2. 对坐标排序后,通过比例划分确定页眉、页脚的临界值
  3. 用动态临界值替换固定阈值,提取正文

完整代码实现

from PyPDF2 import PdfReader
from collections import defaultdict

def get_page_bounds(reader, sample_pages=3):
    y_coords = defaultdict(int)
    max_pages = min(sample_pages, len(reader.pages))
    
    def collect_coords(text, cm, tm, fontDict, fontSize):
        y = round(tm[5], 2)  # 保留两位小数降低精度干扰
        if text.strip():  # 仅统计非空文本的坐标
            y_coords[y] += 1
    
    # 收集样本页面的所有文本y坐标
    for i in range(max_pages):
        page = reader.pages[i]
        page.extract_text(visitor_text=collect_coords)
    
    if not y_coords:
        return (50, 720)  # 无有效坐标时使用默认值兜底
    
    # 排序所有y坐标并计算边界
    sorted_ys = sorted(y_coords.keys())
    total_coords = len(sorted_ys)
    
    # 取底部10%和顶部10%的坐标作为页眉页脚边界,可根据排版调整比例
    footer_cutoff = sorted_ys[int(total_coords * 0.1)]
    header_cutoff = sorted_ys[int(total_coords * 0.9)]
    
    return (footer_cutoff, header_cutoff)

# 实际使用示例
reader = PdfReader("GeoBase_NHNC1_Data_Model_UML_EN.pdf")
footer_bound, header_bound = get_page_bounds(reader)

parts = []
def visitor_body(text, cm, tm, fontDict, fontSize):
    y = tm[5]
    if y > footer_bound and y < header_bound:
        parts.append(text)

# 提取目标页面正文
page = reader.pages[3]
page.extract_text(visitor_text=visitor_body)
text_body = "".join(parts)
print(text_body)

代码说明

  • get_page_bounds函数:通过样本页面收集坐标,统计后动态计算正文边界,默认取前3页避免单页异常
  • 对y坐标做四舍五入处理,减少PDF排版精度差异带来的干扰
  • 按比例划分边界的逻辑可灵活调整,适配不同PDF的页眉页脚占比
  • 保留兜底逻辑,确保无法收集坐标时仍能使用原固定值正常运行

内容的提问来源于stack exchange,提问作者Tara

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 12:07:38