基于PyMuPDF从PDF中提取图片关联的指定文本的方法
基于PyMuPDF从PDF中提取图片关联的指定文本的方法
你好!从你的描述来看,你已经顺利用PyMuPDF提取出了PDF里的图片,现在想要精准提取对应Fig. 6.1的完整图注文本(也就是Fig. 6.1 Insect bites. Linear pruritic papules with central crusts demonstrating the “breakfast, lunch, and dinner” sign. Courtesy Antonio Torrelo, MD.这段内容)对吧?我来给你梳理几个可行的解决思路和代码示例。
首先先修正你之前提取图片代码里的一处小细节:原来的xref = img[page_index]应该改成xref = img[0],因为图片元组的第一个元素才是图片的XREF索引。修正后的提取图片代码如下:
import pymupdf doc = pymupdf.open('sample.pdf') page = doc[0] # 获取目标页面 image_list = page.get_images() page_index = 0 for image_index, img in enumerate(image_list): xref = img[0] # 正确获取图片的XREF pix = pymupdf.Pixmap(doc, xref) if pix.n - pix.alpha > 3: # 处理CMYK格式图片,转成RGB pix = pymupdf.Pixmap(pymupdf.csRGB, pix) pix.save(f"page_{page_index}-image_{image_index}.png")
接下来是核心的“图片关联文本提取”部分,核心思路是利用图片的位置边界框(BBox)匹配附近的文本块,或者直接定位Fig. 6.1文本并提取后续连续内容,以下是两种具体实现方法:
方法一:通过图片位置匹配对应图注
通常PDF里的图注会紧邻图片(上方或下方),我们可以先获取图片的页面坐标范围,再筛选出该范围内的文本块,定位到目标图注:
import pymupdf doc = pymupdf.open('sample.pdf') page = doc[0] # --- 步骤1:提取图片并记录目标图片的边界框 --- image_list = page.get_images() target_img_bbox = None # 存储Fig.6.1对应图片的边界框 for image_index, img in enumerate(image_list): xref = img[0] pix = pymupdf.Pixmap(doc, xref) # 获取当前图片的边界框(格式:x0, y0, x1, y1,对应左上角和右下角坐标) current_bbox = img[7] print(f"图片{image_index}的边界框:{current_bbox}") # 保存图片(和之前逻辑一致) if pix.n - pix.alpha > 3: pix = pymupdf.Pixmap(pymupdf.csRGB, pix) pix.save(f"page_0-image_{image_index}.png") # 如果你已经知道哪个是目标图片,可直接赋值,比如假设是第0张图 target_img_bbox = current_bbox # --- 步骤2:提取图片附近的图注文本 --- # 获取页面所有文本块,sort=True确保按阅读顺序排序 text_blocks = page.get_text("blocks", sort=True) target_fig = "Fig. 6.1" full_caption = "" found_target = False # 定义图片附近的范围:这里假设图注在图片下方200像素内,可根据PDF调整 img_bottom_y = target_img_bbox[3] caption_y_min = img_bottom_y caption_y_max = img_bottom_y + 200 for block in text_blocks: block_x0, block_y0, block_x1, block_y1, block_text, block_no, block_type = block cleaned_text = block_text.strip() # 如果已经找到目标Fig,就合并后续连续文本块,直到空块或超出范围 if found_target: if not cleaned_text or block_y0 > caption_y_max: break full_caption += " " + cleaned_text continue # 检查当前文本块是否在图片附近,且包含目标Fig编号 if caption_y_min < block_y0 < caption_y_max and target_fig in cleaned_text: found_target = True full_caption = cleaned_text # 清理文本中的多余换行和空格 full_caption = ' '.join(full_caption.split()) print("提取到的图注文本:") print(full_caption)
方法二:直接定位Fig.6.1并提取后续文本
如果你的PDF中图注是连续的段落(没有被无关文本打断),可以直接定位到包含Fig. 6.1的文本块,再合并后续文本块得到完整内容:
import pymupdf doc = pymupdf.open('sample.pdf') page = doc[0] # 获取按阅读顺序排序的文本块 text_blocks = page.get_text("blocks", sort=True) target_fig = "Fig. 6.1" full_caption = "" found_target = False for block in text_blocks: cleaned_text = block[4].strip() # 找到目标Fig后,合并后续连续文本 if found_target: # 假设空文本块代表段落结束,可根据实际情况调整判断逻辑 if not cleaned_text: break full_caption += " " + cleaned_text continue # 检测是否包含目标Fig编号 if target_fig in cleaned_text: found_target = True full_caption = cleaned_text # 清理格式 full_caption = ' '.join(full_caption.split()) print(full_caption)
注意事项
- 文本块排序:一定要用
page.get_text("blocks", sort=True),否则文本块可能会乱序,导致提取的内容不连贯。 - 范围调整:如果图注在图片上方,只需把
caption_y_min和caption_y_max改成图片顶部的y坐标上下范围即可。 - 特殊情况处理:如果PDF的图注被拆分成多个零散文本块,或者有干扰文本,可能需要添加更精确的判断(比如检测文本是否包含“Courtesy”这类常见的图注结尾词)。
备注:内容来源于stack exchange,提问作者NPatel
相关产品推荐
相关产品推荐

