You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PyMuPDF提取PDF图片输出损坏,寻求解决方法

PDF图片提取损坏问题解决办法

我使用以下代码提取PDF中的图片,但提取出的图片出现损坏,与预期不符,怀疑PDF中的图片存在加密/特殊编码情况,求可行解决办法:

from PIL import Image


def extract_images_from_pdfs(pdf_list):
    import fitz  # PyMuPDF
    
    output_dir = "C:/path_to_image"
    os.makedirs(output_dir, exist_ok=True)
    
    for pdf_path in pdf_list:
        pdf_name = os.path.splitext(os.path.basename(pdf_path))[0]
       
        # Open the PDF
        pdf_document = fitz.open(pdf_path)
        
        # Track the count of images extracted per page
        image_count = 0
        
        for page_num, page in enumerate(pdf_document):
            # Get the images on this page
            image_list = page.get_images(full=True)
            
            if not image_list:
                print(f"No images found on page {page_num+1} of {pdf_name}")
                continue
            
            # Process each image
            for img_index, img in enumerate(image_list):
                xref = img[0]
                base_image = pdf_document.extract_image(xref)
                
                if base_image:
                    image_bytes = base_image["image"]
                    image_ext = base_image["ext"]
                    
                    # Convert bytes to image
                    image = Image.open(io.BytesIO(image_bytes))
                    
                    # Save the image
                    image_name = f"{pdf_name}_image_{image_count}.{image_ext}"
                    image_path = os.path.join(output_dir, image_name)
                    
                    image.save(image_path)
                    
                    image_count += 1
        
        pdf_document.close()
        print(f"Extracted {image_count} images from {pdf_name}")

问题原因

这类图像损坏通常是因为PDF中的图片带有软掩码(Soft Mask)、透明通道或使用了非标准压缩/加密方式,PyMuPDF直接提取原始图像字节时,未正确处理这些附加属性,导致图像显示异常。

可行解决方案

1. 页面渲染提取(最简单有效)

直接渲染PDF页面为高分辨率位图,绕过原始图像的编码问题,保证内容正确:

import fitz
import os
from PIL import Image

def extract_images_via_rendering(pdf_list):
    output_dir = "C:/path_to_image"
    os.makedirs(output_dir, exist_ok=True)
    
    for pdf_path in pdf_list:
        pdf_name = os.path.splitext(os.path.basename(pdf_path))[0]
        pdf_document = fitz.open(pdf_path)
        image_count = 0
        
        for page_num, page in enumerate(pdf_document):
            # 调整dpi控制清晰度,300足以满足大部分场景
            pix = page.get_pixmap(dpi=300)
            img = Image.frombytes("RGB", [pix.width, pix.height], pix.samples)
            image_name = f"{pdf_name}_page_{page_num+1}.png"
            image_path = os.path.join(output_dir, image_name)
            img.save(image_path)
            image_count += 1
        
        pdf_document.close()
        print(f"通过渲染提取了{image_count}张图片(来自{pdf_name})")

2. 修复原始提取逻辑,处理软掩码

如果需要保留原始图像格式(而非位图),可以尝试检测并合并图像的软掩码:

import fitz
import os
import io
from PIL import Image

def extract_images_with_mask(pdf_list):
    output_dir = "C:/path_to_image"
    os.makedirs(output_dir, exist_ok=True)
    
    for pdf_path in pdf_list:
        pdf_name = os.path.splitext(os.path.basename(pdf_path))[0]
        pdf_document = fitz.open(pdf_path)
        image_count = 0
        
        for page_num, page in enumerate(pdf_document):
            image_list = page.get_images(full=True)
            if not image_list:
                print(f"{pdf_name}的第{page_num+1}页无图片")
                continue
            
            for img_index, img in enumerate(image_list):
                xref = img[0]
                base_image = pdf_document.extract_image(xref)
                if not base_image:
                    continue
                
                image_bytes = base_image["image"]
                image_ext = base_image["ext"]
                img = Image.open(io.BytesIO(image_bytes))
                
                # 检查并处理软掩码
                smask_xref = img[1] if len(img) > 1 else None
                if smask_xref:
                    smask_image = pdf_document.extract_image(smask_xref)
                    if smask_image:
                        smask_bytes = smask_image["image"]
                        smask = Image.open(io.BytesIO(smask_bytes))
                        img.putalpha(smask)
                
                image_name = f"{pdf_name}_image_{image_count}.{image_ext}"
                image_path = os.path.join(output_dir, image_name)
                img.save(image_path)
                image_count += 1
        
        pdf_document.close()
        print(f"提取并修复了{image_count}张图片(来自{pdf_name})")

3. 先解密PDF再提取

如果PDF确实存在加密,可先用qpdf解密后再处理:

  1. 安装qpdf:pip install qpdf(或通过系统包管理器安装)
  2. 执行解密命令:qpdf --decrypt input.pdf output_decrypted.pdf
  3. 使用你原有的代码处理解密后的PDF

内容的提问来源于stack exchange,提问作者ShinyZack123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.13 14:39:52