You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyMuPDF的page.get_text("dict")获取图片xref时出现bad xref错误

解决PyMuPDF提取图片时bad xref错误的实用方案

1. 别直接用get_text返回的xref,先找对真正的图片流引用

page.get_text("dict")返回的图片块里的xref,很多时候指向的是图片的属性字典,不是真正存储图片数据的流对象。你可以换个方式直接拿页面图片的有效xref:

import fitz

def extract_pdf_content(doc):
    img_counter = 0
    for page in doc:
        text_data = page.get_text("dict")
        for block in text_data["blocks"]:
            if block["type"] == 1:  # 识别图片块
                img_counter +=1
                # 直接获取页面所有图片的有效xref列表
                page_imgs = page.get_images(full=True)
                for img_info in page_imgs:
                    valid_xref = img_info[0]
                    try:
                        img_data = doc.extract_image(valid_xref)
                        # 保存图片到本地
                        with open(f"image_{img_counter}.png", "wb") as f:
                            f.write(img_data["image"])
                        # 插入约定的图片标签
                        print(f"<<<image_{img_counter}>>>")
                        break  # 匹配到对应图片就跳出循环
                    except ValueError:
                        continue
            else:  # 文本块直接输出内容
                print(block["text"])

2. 先修复PDF的xref表

如果PDF本身的索引表损坏,直接处理必然报错,先修复再操作:

# 打开原PDF并修复保存,garbage=4会彻底清理无效对象和坏xref
doc = fitz.open("your_pdf.pdf")
doc.save("fixed_pdf.pdf", garbage=4, deflate=True)
doc.close()

# 用修复后的文件执行后续提取逻辑
doc = fitz.open("fixed_pdf.pdf")
# 插入上面的提取代码

3. 换个思路:先拿所有图片,再按位置插标签

放弃从文本字典里拿xref,直接用page.get_images()获取页面所有有效图片,再根据文本块和图片的位置关系,把标签插到对应文本位置:

def process_pdf_with_pos(doc):
    img_counter = 0
    for page in doc:
        # 获取页面所有图片的位置和有效xref
        page_imgs = page.get_images(full=True)
        img_info_list = []
        for img in page_imgs:
            xref = img[0]
            # 解析图片的位置矩形(从PDF对象里提取MediaBox)
            img_obj = doc.xref_object(xref)
            import re
            box_match = re.search(r"/MediaBox\s+\[(\d+)\s+(\d+)\s+(\d+)\s+(\d+)\]", img_obj)
            if box_match:
                x0, y0, x1, y1 = map(float, box_match.groups())
                img_rect = fitz.Rect(x0, y0, x1, y1)
                img_info_list.append((img_rect, xref))
        
        # 把文本块按垂直位置排序(从上到下)
        text_blocks = sorted(page.get_text("dict")["blocks"], key=lambda b: b["bbox"][1])
        img_idx = 0
        block_count = len(text_blocks)
        
        for i in range(block_count):
            block_rect = fitz.Rect(text_blocks[i]["bbox"])
            # 检查当前文本块之前有没有未处理的图片
            while img_idx < len(img_info_list):
                img_rect, img_xref = img_info_list[img_idx]
                # PDF的y轴向下,图片底部y1小于文本块顶部y0,说明图片在文本上方
                if img_rect.y1 < block_rect.y0:
                    img_counter +=1
                    try:
                        img_data = doc.extract_image(img_xref)
                        with open(f"image_{img_counter}.png", "wb") as f:
                            f.write(img_data["image"])
                        print(f"<<<image_{img_counter}>>>")
                    except ValueError:
                        pass
                    img_idx +=1
                else:
                    break
            # 输出当前文本块内容
            if text_blocks[i]["type"] == 0:
                print(text_blocks[i]["text"])
        
        # 处理页面末尾剩余的图片
        while img_idx < len(img_info_list):
            img_counter +=1
            img_rect, img_xref = img_info_list[img_idx]
            try:
                img_data = doc.extract_image(img_xref)
                with open(f"image_{img_counter}.png", "wb") as f:
                    f.write(img_data["image"])
                print(f"<<<image_{img_counter}>>>")
            except ValueError:
                pass
            img_idx +=1

4. 升级PyMuPDF到最新版

旧版本的PyMuPDF存在get_text("dict")返回错误xref的bug,直接升级即可:

pip install --upgrade pymupdf

内容的提问来源于stack exchange,提问作者Gregory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 21:58:26