You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取PDF中的高亮文本及其关联批注

问题:关联PDF高亮文本与对应批注并导出结构化数据

场景与需求

场景:PDF文档中存在高亮文本(例如“is on sales for $100”),该高亮文本附带批注(例如“replace with $99”)
需求:将批注与其指向的高亮文本关联,并存储为DataFrame、表格或JSON等结构化格式

现有尝试代码

import pymupdf

doc = pymupdf.open("demo.pdf")
print("Number of pages:", doc.page_count)  # Check the number of pages

for i in range(doc.page_count):
    page = doc[i]
    annotations = list(page.annots())  # Convert the generator to a list
    if annotations:  # Check if there are any annotations
        print(f"Page {i+1} has {len(annotations)} annotations:")
        for annot in annotations:
            print(annot.info)  # Print detailed annotation info
    else:
        print(f"Page {i+1} has no annotations.")

doc.close()  # Good practice to close the document

运行结果

Number of pages: 12
Page 1 has 2 annotations:
{'content': 'replace with excitment', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''}
{'content': 'Remove "rapid progress" from this sentence', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''}
Page 2 has no annotations.
Page 3 has no annotations.
Page 4 has 1 annotations:
{'content': 'replace with initial idea generation', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''}
Page 5 has no annotations.
Page 6 has no annotations.
Page 7 has no annotations.
Page 8 has no annotations.
Page 9 has no annotations.
Page 10 has no annotations.
Page 11 has no annotations.
Page 12 has no annotations.

当前问题

现有代码仅能输出批注内容,无法获取批注所指向的高亮文本,无法实现二者的关联存储。

解决方案

利用PyMuPDF的批注坐标信息,提取高亮区域对应的文本,再与批注内容关联并导出为结构化数据:

import pymupdf
import pandas as pd

doc = pymupdf.open("demo.pdf")
# 存储关联数据的列表
annot_records = []

for page_idx in range(doc.page_count):
    page = doc[page_idx]
    annotations = page.annots()
    if not annotations:
        continue
    for annot in annotations:
        # 仅处理高亮类型的批注(PyMuPDF中类型代码8对应Highlight)
        if annot.type[0] == 8:
            # 获取批注内容
            comment = annot.info.get("content", "").strip()
            # 提取高亮区域的文本:高亮可能由多个四边形组成,遍历每个区域
            highlighted_text = ""
            for quad_points in annot.quadpoints:
                # 将四边形坐标转换为矩形框
                rect = pymupdf.Quad(quad_points).rect
                # 提取矩形内的文本并拼接
                segment_text = page.get_textbox(rect).strip()
                highlighted_text += f"{segment_text} "
            highlighted_text = highlighted_text.strip()
            # 将数据存入列表
            annot_records.append({
                "页码": page_idx + 1,
                "高亮文本": highlighted_text,
                "批注内容": comment
            })

doc.close()

# 转换为DataFrame并输出
df = pd.DataFrame(annot_records)
print("关联结果:")
print(df)

# 导出为JSON文件
df.to_json("pdf_annotations.json", orient="records", force_ascii=False, indent=4)
print("已导出为pdf_annotations.json")

关键说明

  • 批注类型判断:通过annot.type[0] == 8筛选高亮批注,PyMuPDF中不同批注类型对应不同代码,可通过print(annot.type)查看具体类型信息
  • 高亮文本提取:利用批注的quadpoints属性获取每个高亮区域的坐标(支持跨多行的高亮),再通过page.get_textbox()提取对应区域的文本
  • 结构化存储:将关联数据存入列表后,可直接转换为DataFrame用于分析,或导出为JSON文件持久化存储

内容的提问来源于stack exchange,提问作者jl02

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 06:17:04