如何用Python提取PDF中的高亮文本及其关联批注
问题:关联PDF高亮文本与对应批注并导出结构化数据
场景与需求
场景:PDF文档中存在高亮文本(例如“is on sales for $100”),该高亮文本附带批注(例如“replace with $99”)
需求:将批注与其指向的高亮文本关联,并存储为DataFrame、表格或JSON等结构化格式
现有尝试代码
import pymupdf doc = pymupdf.open("demo.pdf") print("Number of pages:", doc.page_count) # Check the number of pages for i in range(doc.page_count): page = doc[i] annotations = list(page.annots()) # Convert the generator to a list if annotations: # Check if there are any annotations print(f"Page {i+1} has {len(annotations)} annotations:") for annot in annotations: print(annot.info) # Print detailed annotation info else: print(f"Page {i+1} has no annotations.") doc.close() # Good practice to close the document
运行结果
Number of pages: 12 Page 1 has 2 annotations: {'content': 'replace with excitment', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''} {'content': 'Remove "rapid progress" from this sentence', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''} Page 2 has no annotations. Page 3 has no annotations. Page 4 has 1 annotations: {'content': 'replace with initial idea generation', 'name': '', 'title': '', 'creationDate': '', 'modDate': '', 'subject': '', 'id': ''} Page 5 has no annotations. Page 6 has no annotations. Page 7 has no annotations. Page 8 has no annotations. Page 9 has no annotations. Page 10 has no annotations. Page 11 has no annotations. Page 12 has no annotations.
当前问题
现有代码仅能输出批注内容,无法获取批注所指向的高亮文本,无法实现二者的关联存储。
解决方案
利用PyMuPDF的批注坐标信息,提取高亮区域对应的文本,再与批注内容关联并导出为结构化数据:
import pymupdf import pandas as pd doc = pymupdf.open("demo.pdf") # 存储关联数据的列表 annot_records = [] for page_idx in range(doc.page_count): page = doc[page_idx] annotations = page.annots() if not annotations: continue for annot in annotations: # 仅处理高亮类型的批注(PyMuPDF中类型代码8对应Highlight) if annot.type[0] == 8: # 获取批注内容 comment = annot.info.get("content", "").strip() # 提取高亮区域的文本:高亮可能由多个四边形组成,遍历每个区域 highlighted_text = "" for quad_points in annot.quadpoints: # 将四边形坐标转换为矩形框 rect = pymupdf.Quad(quad_points).rect # 提取矩形内的文本并拼接 segment_text = page.get_textbox(rect).strip() highlighted_text += f"{segment_text} " highlighted_text = highlighted_text.strip() # 将数据存入列表 annot_records.append({ "页码": page_idx + 1, "高亮文本": highlighted_text, "批注内容": comment }) doc.close() # 转换为DataFrame并输出 df = pd.DataFrame(annot_records) print("关联结果:") print(df) # 导出为JSON文件 df.to_json("pdf_annotations.json", orient="records", force_ascii=False, indent=4) print("已导出为pdf_annotations.json")
关键说明
- 批注类型判断:通过
annot.type[0] == 8筛选高亮批注,PyMuPDF中不同批注类型对应不同代码,可通过print(annot.type)查看具体类型信息 - 高亮文本提取:利用批注的
quadpoints属性获取每个高亮区域的坐标(支持跨多行的高亮),再通过page.get_textbox()提取对应区域的文本 - 结构化存储:将关联数据存入列表后,可直接转换为DataFrame用于分析,或导出为JSON文件持久化存储
内容的提问来源于stack exchange,提问作者jl02
相关产品推荐
相关产品推荐

