You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

报错OSError: cannot write mode PA as PNG,如何从PDF提取图片?

解决OSError: cannot write mode PA as PNG问题并提取PDF图片

一、报错原因

PA是PyMuPDF(fitz)中带有预乘alpha通道的像素格式,直接用常规方式保存为PNG时会因格式不兼容触发该错误。另外,你的当前代码仅打印页面的pixmap对象,并未真正执行提取或保存图片的操作,这可能偏离了你的需求。

二、解决方案分两种场景:

场景1:提取整个页面为图片(页面截图)

如果需要将PDF页面转为图片,修改代码如下,使用PyMuPDF自带的方法处理格式兼容问题:

import fitz

pdf_file = fitz.open(r"C:\Users\user\Downloads\example.pdf")
for page_index in range(len(pdf_file)):
    page = pdf_file[page_index]
    # 获取页面像素快照,直接用自带save方法自动处理格式
    pix = page.get_pixmap()
    pix.save(f"page_{page_index+1}.png")
    
    # 若需用PIL额外处理,可先转换格式
    # from PIL import Image
    # img = Image.frombytes("RGBA", [pix.width, pix.height], pix.samples)
    # img.save(f"page_{page_index+1}.png")

场景2:提取PDF中嵌入的独立图片

如果要提取PDF里的原生嵌入图片(而非整个页面截图),使用以下代码:

import fitz

pdf_file = fitz.open(r"C:\Users\user\Downloads\example.pdf")
img_count = 0

for page_index in range(len(pdf_file)):
    page = pdf_file[page_index]
    # 获取页面中所有嵌入图片的完整信息
    image_list = page.get_images(full=True)
    
    for img in image_list:
        img_count += 1
        xref = img[0]  # 图片的交叉引用编号
        # 提取图片原始数据和格式信息
        base_image = pdf_file.extract_image(xref)
        image_bytes = base_image["image"]
        img_ext = base_image["ext"]
        
        # 保存提取的图片
        with open(f"extracted_img_{img_count}.{img_ext}", "wb") as f:
            f.write(image_bytes)

补充说明

  • 若提取的图片仍存在格式问题,可通过PIL转换模式解决:
from PIL import Image
import io

img = Image.open(io.BytesIO(image_bytes))
if img.mode == "PA":
    img = img.convert("RGBA")
img.save(f"converted_img_{img_count}.png")

内容的提问来源于stack exchange,提问作者YAŞAR EMRE DOĞRU

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 21:15:41