You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python提取PDF每页标题及高亮文本

问题原因

PyPDF2自带的extractText()方法仅能按读取顺序提取页面纯文本,无法获取字体字号、文本位置、标注等格式元数据,因此既无法区分标题与正文,也无法提取高亮标注内容,只能输出全页文本属于该库本身的能力局限。


无统一字号规则的单页标题提取方案

换用支持读取文本格式属性的PyMuPDF(导入名为fitz)实现,不需要预设全局标题字号,基于单页内的文本特征自动识别,核心识别规则:

  • 标题默认出现在页面顶部1/4区域内,先排除页面中下部的正文内容
  • 单页内字号越大、字体加粗的文本,为标题的概率越高(不依赖全局统一字号,仅在当前页做相对比较)
  • 过滤无效候选:长度超过50字的长文本判定为正文段落,字号小于页面正文平均字号的短文本判定为页眉、页码

代码中的顶部区域占比、标题最大长度阈值可根据自身文档格式调整。
首先安装依赖:
pip install pymupdf

实现代码:

import fitz

def extract_page_titles(pdf_path):
    doc = fitz.open(pdf_path)
    page_titles = []
    for page in doc:
        page_height = page.rect.height
        text_blocks = page.get_text("dict")["blocks"]
        candidate_blocks = []
        # 统计当前页正文基准字号
        font_sizes = []
        for block in text_blocks:
            if "lines" not in block:
                continue
            for line in block["lines"]:
                for span in line["spans"]:
                    if span["text"].strip():
                        font_sizes.append(span["size"])
        base_font_size = sum(font_sizes)/len(font_sizes) if font_sizes else 10
        # 筛选顶部区域的候选标题
        for block in text_blocks:
            if "lines" not in block or block["bbox"][1] > page_height/4:
                continue
            block_text = ""
            max_size = 0
            is_bold = False
            for line in block["lines"]:
                for span in line["spans"]:
                    block_text += span["text"]
                    if span["size"] > max_size:
                        max_size = span["size"]
                    if "bold" in span["font"].lower():
                        is_bold = True
            block_text = block_text.strip()
            # 过滤不符合规则的候选
            if not block_text or len(block_text) > 50 or max_size < base_font_size:
                continue
            # 计算匹配权重
            weight = max_size * (2 if is_bold else 1)
            candidate_blocks.append((weight, block_text))
        # 取权重最高的作为当前页标题
        if candidate_blocks:
            candidate_blocks.sort(reverse=True, key=lambda x:x[0])
            page_titles.append(candidate_blocks[0][1])
        else:
            page_titles.append("未识别到标题")
    doc.close()
    return page_titles

# 调用示例
titles = extract_page_titles("Test2.pdf")
for idx, title in enumerate(titles, 1):
    print(f"第{idx}页标题:{title}")

PDF高亮文本提取方法

同样基于PyMuPDF实现,高亮属于PDF的内置注释类型,直接遍历页面注释筛选高亮类注释,提取关联区域的文本即可,实现代码:

import fitz

def extract_highlight_text(pdf_path):
    doc = fitz.open(pdf_path)
    all_highlights = []
    for page_num, page in enumerate(doc, 1):
        page_highlights = []
        annot = page.first_annot
        while annot:
            # 8对应PyMuPDF中定义的高亮注释类型
            if annot.type[0] == 8:
                highlight_text = page.get_textbox(annot.rect).strip()
                if highlight_text:
                    page_highlights.append(highlight_text)
            annot = annot.next
        all_highlights.append({
            "page": page_num,
            "highlights": page_highlights
        })
    doc.close()
    return all_highlights

# 调用示例
highlight_result = extract_highlight_text("Test2.pdf")
for item in highlight_result:
    if item["highlights"]:
        print(f"第{item['page']}页高亮内容:")
        for h in item["highlights"]:
            print(f"- {h}")

注意:如果待处理PDF是扫描生成的图片版PDF,以上方法无法直接识别文本和标题,需要先做OCR处理后再匹配规则。

内容的提问来源于stack exchange,提问作者Prajkta Mangulkar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 11:15:42