使用PyMuPDF提取PDF图片不全的问题求助
PDF图片提取不全问题排查与解决
问题描述
尝试从设备订单发票PDF中提取图片,但每次运行代码仅能提取每页8-9张图片中的4张。即使使用他人编写的图片计数代码,结果依然相同,怀疑部分PDF与PyMuPDF的部分函数存在兼容性问题。
原代码如下:
def extract_images(model_nums, file): image_num = 0 doc = fitz.open(file) # new directories that will hold images all_path = os.path.join(os.getcwd(), "All Files") if not os.path.exists('All Files'): os.mkdir(all_path) if not os.path.exists(sport_id): os.mkdir(sport_path) for i in range(doc.page_count): print("Page: "+ str(i)) images = doc.get_page_images(i) for img in images: xref = img[0] pix = fitz.Pixmap(doc, xref) pix.save(f"{all_path}/{model_nums[image_num]}.jpg") pix = None image_num += 1
问题原因
PyMuPDF的get_page_images()方法默认仅提取页面内容流中直接引用的图像,但PDF中的图像可能通过以下方式隐藏,导致无法被捕获:
- 图像嵌入在表单XObject(可复用的页面模板对象)中
- 图像被压缩在页面资源字典但未直接在内容流中引用
- 同一图像被多个页面复用,原代码未处理去重但实际可能漏提其他未引用的实例
- 部分图像采用特殊编码或加密格式
解决方案
修改代码,遍历PDF中所有对象,检查并提取所有图像资源,同时处理重复图像:
import os import fitz def extract_images(model_nums, file): image_num = 0 processed_xrefs = set() # 记录已处理的图像xref,避免重复提取 doc = fitz.open(file) # 创建存储目录 all_path = os.path.join(os.getcwd(), "All Files") os.makedirs(all_path, exist_ok=True) # 注意:原代码中sport_id和sport_path未定义,需补充定义或移除相关逻辑 # sport_path = os.path.join(os.getcwd(), sport_id) # os.makedirs(sport_path, exist_ok=True) # 遍历PDF中所有对象,查找图像 for xref in range(1, doc.xref_length()): # 检查当前对象是否为图像 if doc.xref_get_key(xref, "Subtype")[1] == "/Image": if xref in processed_xrefs: continue processed_xrefs.add(xref) try: pix = fitz.Pixmap(doc, xref) # 处理CMYK颜色空间的图像,转换为RGB if pix.colorspace == fitz.csCMYK: pix = fitz.Pixmap(fitz.csRGB, pix) # 保存图像,注意model_nums长度需匹配图像总数 if image_num < len(model_nums): pix.save(f"{all_path}/{model_nums[image_num]}.jpg") else: # 如果model_nums不够,用默认命名 pix.save(f"{all_path}/image_{xref}.jpg") pix = None image_num += 1 except Exception as e: print(f"提取图像xref {xref}失败: {str(e)}") continue print(f"共提取{image_num}张图像") doc.close()
关键改进点
- 遍历PDF所有对象而非仅页面引用的图像,覆盖所有嵌入的图像资源
- 用
processed_xrefs集合避免重复提取同一图像 - 增加CMYK转RGB的处理,避免部分图像保存失败
- 增加异常捕获,防止单个图像提取失败导致程序中断
- 使用
os.makedirs(exist_ok=True)简化目录创建逻辑
内容的提问来源于stack exchange,提问作者Asia Vassos
相关产品推荐
相关产品推荐

