You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab使用pdf2image转PDF为图片报UnidentifiedImageError错误

问题场景

使用pdf2image库调用convert_from_path方法实现PDF转图片功能,相同代码在本地Jupyter Notebook环境可正常运行,迁移到Google Colab环境后持续报错,已提前安装poppler-utils依赖。

运行代码如下:

import pytesseract
import shutil
import os
import random
try:
 from PIL import Image
except ImportError:
 import Image

import cv2
from pdf2image import convert_from_path

from google.colab import files
uploaded = files.upload()

pages=convert_from_path(r"foo.pdf",500)

报错信息

UnidentifiedImageError: cannot identify image file <_io.BytesIO object at 0x7fe6abfae5f0>

报错中文释义:无法识别图像文件错误:无法识别对应内存字节流的图像内容

故障原因
  1. Google Colab环境中poppler-utils安装后,默认二进制路径不在pdf2image的自动检索范围内,库调用poppler解析PDF时生成无效字节流,触发PIL的图像识别异常
  2. files.upload()上传文件后,可能存在文件名拼写偏差、文件未实际存入当前工作目录的路径匹配问题
解决方法

按以下步骤操作即可修复:

  • 执行命令重新确认poppler-utils安装完整:!apt-get install -y poppler-utils
  • 执行!ls命令查看当前工作目录文件列表,确认上传的PDF文件名为foo.pdf,无拼写错误、无后缀名隐藏问题
  • 调用转换方法时显式指定Colab环境下的poppler路径,Colab中poppler可执行文件默认存放在/usr/bin目录
  • 若仍存在路径匹配问题,可直接读取上传接口返回的文件字节流,用convert_from_bytes方法完成转换,绕过文件路径校验逻辑

修复后的可运行代码:

# 安装依赖
!apt-get install -y poppler-utils
!pip install -q pdf2image pillow opencv-python pytesseract

import os
from PIL import Image
import cv2
from pdf2image import convert_from_path, convert_from_bytes
from google.colab import files

# 上传PDF文件
uploaded = files.upload()

# 校验文件存在
assert "foo.pdf" in os.listdir(), "foo.pdf 未在当前工作目录找到,请检查上传文件"

# 方式1:路径读取,显式指定poppler路径
pages = convert_from_path(
    pdf_path="foo.pdf",
    dpi=500,
    poppler_path="/usr/bin"
)

# 若方式1仍报错,使用方式2:直接读取上传字节流转换
# pages = convert_from_bytes(uploaded["foo.pdf"], dpi=500, poppler_path="/usr/bin")

内容的提问来源于stack exchange,提问作者Ace Purohit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 18:31:23