如何使用PyMuPDF匹配提取出的PDF链接与文本?
使用PyMuPDF匹配PDF文本与链接的最佳方法
PyMuPDF 中 page.get_links() 返回的每个链接对象都包含**边界框(bbox)**信息,这是匹配对应文本的核心依据。下面是两种实用的实现方案:
方案一:精准匹配文本块(适合复杂排版)
通过 page.get_text("dict") 提取带位置坐标的文本结构,再逐一比对链接边界框与文本块的重叠关系:
import fitz # PyMuPDF def match_links_and_text(page): # 提取带位置信息的文本字典(包含块、行、字符层级的坐标) text_struct = page.get_text("dict") links = page.get_links() matched_pairs = [] for link in links: link_bbox = fitz.Rect(link["from"]) linked_text = "" # 遍历所有文本块 for block in text_struct["blocks"]: if block["type"] != 0: # 跳过非文本块(如图片) continue for line in block["lines"]: for span in line["spans"]: span_bbox = fitz.Rect(span["bbox"]) # 判断文本段与链接框是否重叠(可根据需求调整匹配逻辑) if span_bbox.intersects(link_bbox): linked_text += span["text"] matched_pairs.append({ "url": link.get("uri"), "text": linked_text.strip(), "bbox": link["from"] }) return matched_pairs # 调用示例 doc = fitz.open("target.pdf") page = doc[0] results = match_links_and_text(page) for pair in results: print(f"关联文本: {pair['text']}\n对应链接: {pair['url']}\n")
方案二:直接提取链接区域文本(简洁高效)
利用 page.get_textbox(bbox) 方法,直接传入链接的边界框,提取该区域内的文本,适合常规排版的PDF:
import fitz def match_links_and_text_simple(page): links = page.get_links() matched_pairs = [] for link in links: link_bbox = fitz.Rect(link["from"]) # 提取链接框覆盖的文本 linked_text = page.get_textbox(link_bbox).strip() matched_pairs.append({ "url": link.get("uri"), "text": linked_text, "bbox": link["from"] }) return matched_pairs # 调用示例 doc = fitz.open("target.pdf") page = doc[0] results = match_links_and_text_simple(page) for pair in results: print(f"关联文本: {pair['text']}\n对应链接: {pair['url']}\n")
注意事项
- 若链接框与文本框不完全对齐,可将
intersects替换为contains,或自定义重叠比例阈值来优化匹配精度。 - 对于内部跳转链接(
link["kind"] == fitz.LINK_GOTO),可通过link["page"]link["dest"]定位目标页面,匹配逻辑与外部链接一致。 - 多列、嵌套排版的PDF建议使用方案一,能更精准捕获每个文本片段。
内容的提问来源于stack exchange,提问作者Cai
相关产品推荐
相关产品推荐

