You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pdf2image的convert_from_path时遇BytesIO对象无法识别错误

解决pdf2image convert_from_path抛出UnidentifiedImageError的问题

问题场景

使用pdf2image库的convert_from_path方法提取PDF页面时,突然抛出UnidentifiedImageError,代码此前能正常运行。

报错信息

<ipython-input-45-4ebf020b9136> in <cell line: 1>()
      1 for pdf in list_of_pdfs:
----> 2   images = convert_from_path(pdf,first_page= 1,last_page=2)

2 frames
/usr/local/lib/python3.10/dist-packages/pdf2image/pdf2image.py in convert_from_path(pdf_path, dpi, output_folder, first_page, last_page, fmt, jpegopt, thread_count, userpw, ownerpw, use_cropbox, strict, transparent, single_file, output_file, poppler_path, grayscale, size, paths_only, use_pdftocairo, timeout, hide_annotations)
    266                 )
    267             else:
---> 268                 images += parse_buffer_func(data)
    269     finally:
    270         if auto_temp_dir:

/usr/local/lib/python3.10/dist-packages/pdf2image/parsers.py in parse_buffer_to_ppm(data)
     26         size_x, size_y = tuple(size.split(b" "))
     27         file_size = len(code) + len(size) + len(rgb) + 3 + int(size_x) * int(size_y) * 3
---> 28         images.append(Image.open(BytesIO(data[index : index + file_size])))
     29         index += file_size
     30 

/usr/local/lib/python3.10/dist-packages/PIL/Image.py in open(fp, mode, formats)
   3281                 raise
   3282         return None
-> 3283 
   3284     im = _open_core(fp, filename, prefix, formats)
   3285 

UnidentifiedImageError: cannot identify image file <_io.BytesIO object at 0x7820221cd290>

运行代码

import io
from io import BytesIO
from PIL import Image
from pdf2image import convert_from_path

pdf_list = ['path_to_pdf.pdf','path_to_pdf2.pdf']
for pdf in pdf_list:
  images = convert_from_path(pdf,first_page= 1,last_page=2)

排查与解决方法

1. 检查PDF文件完整性

  • 手动打开目标PDF,确认第1-2页能正常显示,排除文件损坏、内容缺失的情况。
  • 若PDF来自外部渠道,重新获取原始文件,避免传输过程中出现的文件损坏。

2. 验证poppler环境

pdf2image依赖poppler工具,环境变动可能引发异常:

  • 在终端执行pdftoppm -v(Linux/macOS)或对应Windows命令,确认poppler能正常输出版本信息。
  • 若最近更新过poppler或系统,尝试回滚到之前的稳定版本;未安装或版本异常时重新安装:
    • Linux:sudo apt-get install poppler-utils
    • macOS:brew install poppler
    • Windows:下载预编译包并配置环境变量,或在convert_from_path中通过poppler_path参数指定poppler的bin目录路径。

3. 调整convert_from_path参数

  • 先移除first_page和last_page参数,尝试提取整个PDF,缩小问题范围。
  • 指定输出格式为jpeg或png,规避默认ppm格式可能的解析问题:
    images = convert_from_path(pdf, first_page=1, last_page=2, fmt='jpeg')
    
  • 增加dpi参数(如dpi=300),提升图像分辨率,解决低分辨率下的解析异常。

4. 回滚依赖库版本

依赖库版本更新可能引入兼容性问题:

  • 执行pip show pdf2image pillow查看当前版本,回滚到之前能正常运行的版本:
    pip install pdf2image==1.16.0 pillow==9.5.0
    

5. 单文件测试

从pdf_list中逐个取出PDF单独测试,定位是特定文件导致的问题,还是批量处理时的异常。若为单个文件问题,可先修复该PDF或转换格式后再处理。

内容的提问来源于stack exchange,提问作者Harshal Naik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 09:16:21