You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Fitz库提取PDF图片报错:'Page'对象无'getImageList'属性

PDF图片提取报错:AttributeError: 'Page' object has no attribute 'getImageList'

我尝试用Fitz库提取PDF中的图片,使用的代码如下:

import fitz 
import io
from PIL import Image


pdf_file = fitz.open("my_file_pdf.pdf")


for page_index in range(len(pdf_file)):
    # get the page itself
    page = pdf_file[page_index]
    image_list = page.getImageList()
    # printing number of images found in this page
    if image_list:
        print(f"[+] Found  {len(image_list)} images in page {page_index}")
    else:
        print("[!] No images found on the given pdf page", page_index)
    for image_index, img in enumerate(page.getImageList(), start=1):
        print(img)
        print(image_index)
        # get the XREF of the image
        xref = img[0]
        # extract the image bytes
        base_image = pdf_file.extractImage(xref)
        image_bytes = base_image["image"]
        # get the image extension
        image_ext = base_image["ext"]
        # load it to PIL
        image = Image.open(io.BytesIO(image_bytes))
        # save it to local disk
        image.save(open(f"image{page_index+1}_{image_index}.{image_ext}", "wb")) 

运行代码时出现如下错误:

AttributeError                            Traceback (most recent call last)
<ipython-input-1-e5b882e88684> in <module>
     11     # get the page itself
     12     page = pdf_file[page_index]
---> 13     image_list = page.getImageList()
     14     # printing number of images found in this page
     15     if image_list:

AttributeError: 'Page' object has no attribute 'getImageList'

根据文档说明该函数使用方式正确,但实际运行报错,问题出在哪里?


问题原因与解决方法

这个报错是PyMuPDF(即fitz库)的版本差异导致的:

  • 旧版本PyMuPDF中,Page.getImageList()是获取页面图片的有效API;
  • 但在PyMuPDF 1.18.0及以后的版本中,该方法被废弃并移除,替换为符合PEP8规范的page.get_images()方法,同时pdf_file.extractImage()也被重命名为pdf_file.extract_image()。

修正后的代码如下:

import fitz 
import io
from PIL import Image

pdf_file = fitz.open("my_file_pdf.pdf")

for page_index in range(len(pdf_file)):
    page = pdf_file[page_index]
    # 替换为新版本的get_images()方法
    image_list = page.get_images()
    if image_list:
        print(f"[+] Found  {len(image_list)} images in page {page_index}")
    else:
        print("[!] No images found on the given pdf page", page_index)
    # 统一使用get_images()
    for image_index, img in enumerate(page.get_images(), start=1):
        xref = img[0]
        # 使用新版本的extract_image()
        base_image = pdf_file.extract_image(xref)
        image_bytes = base_image["image"]
        image_ext = base_image["ext"]
        image = Image.open(io.BytesIO(image_bytes))
        image.save(open(f"image{page_index+1}_{image_index}.{image_ext}", "wb")) 

替换后代码即可正常运行,提取PDF中的图片。

内容的提问来源于stack exchange,提问作者user60005003

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 09:36:09