使用pdf2image的convert_from_path时遇BytesIO对象无法识别错误
解决pdf2image convert_from_path抛出UnidentifiedImageError的问题
问题场景
使用pdf2image库的convert_from_path方法提取PDF页面时,突然抛出UnidentifiedImageError,代码此前能正常运行。
报错信息
<ipython-input-45-4ebf020b9136> in <cell line: 1>() 1 for pdf in list_of_pdfs: ----> 2 images = convert_from_path(pdf,first_page= 1,last_page=2) 2 frames /usr/local/lib/python3.10/dist-packages/pdf2image/pdf2image.py in convert_from_path(pdf_path, dpi, output_folder, first_page, last_page, fmt, jpegopt, thread_count, userpw, ownerpw, use_cropbox, strict, transparent, single_file, output_file, poppler_path, grayscale, size, paths_only, use_pdftocairo, timeout, hide_annotations) 266 ) 267 else: ---> 268 images += parse_buffer_func(data) 269 finally: 270 if auto_temp_dir: /usr/local/lib/python3.10/dist-packages/pdf2image/parsers.py in parse_buffer_to_ppm(data) 26 size_x, size_y = tuple(size.split(b" ")) 27 file_size = len(code) + len(size) + len(rgb) + 3 + int(size_x) * int(size_y) * 3 ---> 28 images.append(Image.open(BytesIO(data[index : index + file_size]))) 29 index += file_size 30 /usr/local/lib/python3.10/dist-packages/PIL/Image.py in open(fp, mode, formats) 3281 raise 3282 return None -> 3283 3284 im = _open_core(fp, filename, prefix, formats) 3285 UnidentifiedImageError: cannot identify image file <_io.BytesIO object at 0x7820221cd290>
运行代码
import io from io import BytesIO from PIL import Image from pdf2image import convert_from_path pdf_list = ['path_to_pdf.pdf','path_to_pdf2.pdf'] for pdf in pdf_list: images = convert_from_path(pdf,first_page= 1,last_page=2)
排查与解决方法
1. 检查PDF文件完整性
- 手动打开目标PDF,确认第1-2页能正常显示,排除文件损坏、内容缺失的情况。
- 若PDF来自外部渠道,重新获取原始文件,避免传输过程中出现的文件损坏。
2. 验证poppler环境
pdf2image依赖poppler工具,环境变动可能引发异常:
- 在终端执行
pdftoppm -v(Linux/macOS)或对应Windows命令,确认poppler能正常输出版本信息。 - 若最近更新过poppler或系统,尝试回滚到之前的稳定版本;未安装或版本异常时重新安装:
- Linux:
sudo apt-get install poppler-utils - macOS:
brew install poppler - Windows:下载预编译包并配置环境变量,或在
convert_from_path中通过poppler_path参数指定poppler的bin目录路径。
- Linux:
3. 调整convert_from_path参数
- 先移除
first_page和last_page参数,尝试提取整个PDF,缩小问题范围。 - 指定输出格式为
jpeg或png,规避默认ppm格式可能的解析问题:images = convert_from_path(pdf, first_page=1, last_page=2, fmt='jpeg') - 增加
dpi参数(如dpi=300),提升图像分辨率,解决低分辨率下的解析异常。
4. 回滚依赖库版本
依赖库版本更新可能引入兼容性问题:
- 执行
pip show pdf2image pillow查看当前版本,回滚到之前能正常运行的版本:pip install pdf2image==1.16.0 pillow==9.5.0
5. 单文件测试
从pdf_list中逐个取出PDF单独测试,定位是特定文件导致的问题,还是批量处理时的异常。若为单个文件问题,可先修复该PDF或转换格式后再处理。
内容的提问来源于stack exchange,提问作者Harshal Naik
相关产品推荐
相关产品推荐

