如何用pyMuPDF读取PDF中流格式嵌入图片的宽高与DPI信息
解决PyMuPDF无法提取PDF嵌入图片及获取尺寸/DPI的问题
核心问题分析
你拿到的stream是PDF的内容流(包含页面绘图指令,比如cm是坐标变换、gs是图形状态设置),不是图片本身的二进制数据,所以直接用PIL读取必然失败。page.get_images()返回空,说明图片不是页面资源字典直接引用的XObject,可能是嵌套在可选内容组(OCG,你的内容流里有/OC /Pr12 BDC标记)、或者是inline图像(直接嵌入内容流的图像数据)。
解决方案
1. 遍历所有PDF对象,查找图片XObject
PDF中的嵌入图片通常以/XObject类型、/Subtype /Image的对象存在,直接遍历所有xref编号筛选这类对象:
import fitz from PIL import Image import io pdf_file = fitz.open(filepath) # 遍历所有xref对象(从1开始,0是根对象) for xref in range(1, pdf_file.xref_length()): if pdf_file.xref_is_stream(xref): # 获取对象的字典内容,判断是否为图片XObject obj_dict = pdf_file.xref_object(xref, compressed=False) if "/Subtype /Image" in obj_dict: # 提取图片基本参数 img_width = int(pdf_file.xref_get_key(xref, "/Width")[1]) img_height = int(pdf_file.xref_get_key(xref, "/Height")[1]) print(f"找到图片:尺寸 {img_width}x{img_height}") # 提取图片二进制流并读取 img_stream = pdf_file.xref_stream(xref) try: img = Image.open(io.BytesIO(img_stream)) print(f"图片格式:{img.format}") # 读取图片自带的DPI(如果是TIFF等格式可能携带) if "dpi" in img.info: print(f"图片自带DPI:{img.info['dpi']}") else: # 通过PDF的BBox计算DPI(1用户单位=1/72英寸) bbox_str = pdf_file.xref_get_key(xref, "/BBox")[1] bbox = eval(bbox_str) # 格式为[llx, lly, urx, ury] bbox_width = bbox[2] - bbox[0] bbox_height = bbox[3] - bbox[1] dpi_x = (img_width * 72) / bbox_width dpi_y = (img_height * 72) / bbox_height print(f"计算得到DPI:{dpi_x:.2f}x{dpi_y:.2f}") except Exception as e: print(f"读取图片失败:{str(e)}")
2. 解析内容流,提取inline图像
如果图片是直接嵌入内容流的inline图像(以BI开头、EI结尾标记),可以通过正则匹配提取:
import fitz import re from PIL import Image import io pdf_file = fitz.open(filepath) page = pdf_file[0] # 获取内容流的原始文本 content_text = page.get_contents_text() # 匹配BI和EI之间的inline图像数据 inline_img_pattern = re.compile(r'BI(.*?)EI', re.DOTALL) matches = inline_img_pattern.findall(content_text) for img_segment in matches: # 找到ID标记(分隔图像参数和二进制数据) id_pos = img_segment.find('ID') if id_pos == -1: continue # 提取十六进制格式的图像数据并转二进制 img_hex = img_segment[id_pos+2:].strip().replace('\n', '').replace('\r', '') try: img_bytes = bytes.fromhex(img_hex) img = Image.open(io.BytesIO(img_bytes)) print(f"Inline图片尺寸:{img.size}") if "dpi" in img.info: print(f"Inline图片DPI:{img.info['dpi']}") except Exception as e: print(f"处理inline图片失败:{str(e)}")
3. 针对OCG图层图片的补充说明
你的内容流里有/OC /Pr12 BDC,说明图片在可选内容组中。上述遍历xref的方法依然能找到对应的图片XObject,因为OCG只是控制显示,不影响对象的存储结构。
内容的提问来源于stack exchange,提问作者Márcio Duarte
相关产品推荐
相关产品推荐

