You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyMuPDF提取PDF图片不全的问题求助

PDF图片提取不全问题排查与解决

问题描述

尝试从设备订单发票PDF中提取图片,但每次运行代码仅能提取每页8-9张图片中的4张。即使使用他人编写的图片计数代码,结果依然相同,怀疑部分PDF与PyMuPDF的部分函数存在兼容性问题。

原代码如下:

def extract_images(model_nums, file):
    image_num = 0

    doc = fitz.open(file)

    # new directories that will hold images
    all_path = os.path.join(os.getcwd(), "All Files")
    if not os.path.exists('All Files'):
        os.mkdir(all_path)
    if not os.path.exists(sport_id):
        os.mkdir(sport_path)

    for i in range(doc.page_count):
        print("Page: "+ str(i))
    
        images = doc.get_page_images(i)

        for img in images:  
            xref = img[0]
            pix = fitz.Pixmap(doc, xref)
            pix.save(f"{all_path}/{model_nums[image_num]}.jpg")
            pix = None

            image_num += 1

问题原因

PyMuPDF的get_page_images()方法默认仅提取页面内容流中直接引用的图像,但PDF中的图像可能通过以下方式隐藏,导致无法被捕获:

  • 图像嵌入在表单XObject(可复用的页面模板对象)中
  • 图像被压缩在页面资源字典但未直接在内容流中引用
  • 同一图像被多个页面复用,原代码未处理去重但实际可能漏提其他未引用的实例
  • 部分图像采用特殊编码或加密格式

解决方案

修改代码,遍历PDF中所有对象,检查并提取所有图像资源,同时处理重复图像:

import os
import fitz

def extract_images(model_nums, file):
    image_num = 0
    processed_xrefs = set()  # 记录已处理的图像xref,避免重复提取

    doc = fitz.open(file)

    # 创建存储目录
    all_path = os.path.join(os.getcwd(), "All Files")
    os.makedirs(all_path, exist_ok=True)
    
    # 注意:原代码中sport_id和sport_path未定义,需补充定义或移除相关逻辑
    # sport_path = os.path.join(os.getcwd(), sport_id)
    # os.makedirs(sport_path, exist_ok=True)

    # 遍历PDF中所有对象,查找图像
    for xref in range(1, doc.xref_length()):
        # 检查当前对象是否为图像
        if doc.xref_get_key(xref, "Subtype")[1] == "/Image":
            if xref in processed_xrefs:
                continue
            processed_xrefs.add(xref)

            try:
                pix = fitz.Pixmap(doc, xref)
                # 处理CMYK颜色空间的图像,转换为RGB
                if pix.colorspace == fitz.csCMYK:
                    pix = fitz.Pixmap(fitz.csRGB, pix)
                
                # 保存图像,注意model_nums长度需匹配图像总数
                if image_num < len(model_nums):
                    pix.save(f"{all_path}/{model_nums[image_num]}.jpg")
                else:
                    # 如果model_nums不够,用默认命名
                    pix.save(f"{all_path}/image_{xref}.jpg")
                
                pix = None
                image_num += 1
            except Exception as e:
                print(f"提取图像xref {xref}失败: {str(e)}")
                continue

    print(f"共提取{image_num}张图像")
    doc.close()

关键改进点

  • 遍历PDF所有对象而非仅页面引用的图像,覆盖所有嵌入的图像资源
  • 用processed_xrefs集合避免重复提取同一图像
  • 增加CMYK转RGB的处理,避免部分图像保存失败
  • 增加异常捕获,防止单个图像提取失败导致程序中断
  • 使用os.makedirs(exist_ok=True)简化目录创建逻辑

内容的提问来源于stack exchange,提问作者Asia Vassos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.10 01:04:52