You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于PyMuPDF从PDF中提取图片关联的指定文本的方法

基于PyMuPDF从PDF中提取图片关联的指定文本的方法

你好!从你的描述来看,你已经顺利用PyMuPDF提取出了PDF里的图片,现在想要精准提取对应Fig. 6.1的完整图注文本(也就是Fig. 6.1 Insect bites. Linear pruritic papules with central crusts demonstrating the “breakfast, lunch, and dinner” sign. Courtesy Antonio Torrelo, MD.这段内容)对吧?我来给你梳理几个可行的解决思路和代码示例。

首先先修正你之前提取图片代码里的一处小细节:原来的xref = img[page_index]应该改成xref = img[0],因为图片元组的第一个元素才是图片的XREF索引。修正后的提取图片代码如下:

import pymupdf

doc = pymupdf.open('sample.pdf')
page = doc[0]  # 获取目标页面

image_list = page.get_images()
page_index = 0

for image_index, img in enumerate(image_list):
    xref = img[0]  # 正确获取图片的XREF
    pix = pymupdf.Pixmap(doc, xref)

    if pix.n - pix.alpha > 3:  # 处理CMYK格式图片,转成RGB
        pix = pymupdf.Pixmap(pymupdf.csRGB, pix)
    pix.save(f"page_{page_index}-image_{image_index}.png")

接下来是核心的“图片关联文本提取”部分,核心思路是利用图片的位置边界框(BBox)匹配附近的文本块,或者直接定位Fig. 6.1文本并提取后续连续内容,以下是两种具体实现方法:

方法一:通过图片位置匹配对应图注

通常PDF里的图注会紧邻图片(上方或下方),我们可以先获取图片的页面坐标范围,再筛选出该范围内的文本块,定位到目标图注:

import pymupdf

doc = pymupdf.open('sample.pdf')
page = doc[0]

# --- 步骤1:提取图片并记录目标图片的边界框 ---
image_list = page.get_images()
target_img_bbox = None  # 存储Fig.6.1对应图片的边界框

for image_index, img in enumerate(image_list):
    xref = img[0]
    pix = pymupdf.Pixmap(doc, xref)
    
    # 获取当前图片的边界框(格式:x0, y0, x1, y1,对应左上角和右下角坐标)
    current_bbox = img[7]
    print(f"图片{image_index}的边界框:{current_bbox}")
    
    # 保存图片(和之前逻辑一致)
    if pix.n - pix.alpha > 3:
        pix = pymupdf.Pixmap(pymupdf.csRGB, pix)
    pix.save(f"page_0-image_{image_index}.png")
    
    # 如果你已经知道哪个是目标图片,可直接赋值,比如假设是第0张图
    target_img_bbox = current_bbox

# --- 步骤2:提取图片附近的图注文本 ---
# 获取页面所有文本块,sort=True确保按阅读顺序排序
text_blocks = page.get_text("blocks", sort=True)
target_fig = "Fig. 6.1"
full_caption = ""
found_target = False

# 定义图片附近的范围:这里假设图注在图片下方200像素内,可根据PDF调整
img_bottom_y = target_img_bbox[3]
caption_y_min = img_bottom_y
caption_y_max = img_bottom_y + 200

for block in text_blocks:
    block_x0, block_y0, block_x1, block_y1, block_text, block_no, block_type = block
    cleaned_text = block_text.strip()
    
    # 如果已经找到目标Fig,就合并后续连续文本块,直到空块或超出范围
    if found_target:
        if not cleaned_text or block_y0 > caption_y_max:
            break
        full_caption += " " + cleaned_text
        continue
    
    # 检查当前文本块是否在图片附近,且包含目标Fig编号
    if caption_y_min < block_y0 < caption_y_max and target_fig in cleaned_text:
        found_target = True
        full_caption = cleaned_text

# 清理文本中的多余换行和空格
full_caption = ' '.join(full_caption.split())
print("提取到的图注文本:")
print(full_caption)

方法二:直接定位Fig.6.1并提取后续文本

如果你的PDF中图注是连续的段落(没有被无关文本打断),可以直接定位到包含Fig. 6.1的文本块,再合并后续文本块得到完整内容:

import pymupdf

doc = pymupdf.open('sample.pdf')
page = doc[0]

# 获取按阅读顺序排序的文本块
text_blocks = page.get_text("blocks", sort=True)
target_fig = "Fig. 6.1"
full_caption = ""
found_target = False

for block in text_blocks:
    cleaned_text = block[4].strip()
    
    # 找到目标Fig后,合并后续连续文本
    if found_target:
        # 假设空文本块代表段落结束,可根据实际情况调整判断逻辑
        if not cleaned_text:
            break
        full_caption += " " + cleaned_text
        continue
    
    # 检测是否包含目标Fig编号
    if target_fig in cleaned_text:
        found_target = True
        full_caption = cleaned_text

# 清理格式
full_caption = ' '.join(full_caption.split())
print(full_caption)

注意事项

  1. 文本块排序:一定要用page.get_text("blocks", sort=True),否则文本块可能会乱序,导致提取的内容不连贯。
  2. 范围调整:如果图注在图片上方,只需把caption_y_min和caption_y_max改成图片顶部的y坐标上下范围即可。
  3. 特殊情况处理:如果PDF的图注被拆分成多个零散文本块,或者有干扰文本,可能需要添加更精确的判断(比如检测文本是否包含“Courtesy”这类常见的图注结尾词)。

备注:内容来源于stack exchange,提问作者NPatel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.14 14:18:00