You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python PIL读取单页TIFF/PDF报UnidentifiedImageError如何解决

问题原因

你遇到的PIL.UnidentifiedImageError报错由两个独立的兼容性问题导致:

  1. 你的Pillow库编译时没有关联libtiff依赖,无法识别TIFF格式的文件结构
  2. PIL本身不支持直接读取PDF格式文件,需要额外工具链做格式转换
修复步骤

第一步:安装系统依赖(以你使用的macOS环境为例)

  • 安装TIFF解析依赖:brew install libtiff
  • 安装PDF解析依赖:brew install poppler

第二步:重装Pillow保证依赖生效,安装必要的Python包

pip uninstall -y pillow
pip install pillow --no-cache-dir
pip install pdf2image
调整后的兼容代码
import pytesseract
from PIL import Image
from pdf2image import convert_from_path
import os

def ocr_process(image_obj, base_output_name, lang='fra'):
    # 提取文本
    text = pytesseract.image_to_string(image_obj, lang=lang) + '\n\n\n\n'
    with open(f'{base_output_name}.txt', 'w') as fp: 
        fp.write(text)
    # 提取边界框
    text = pytesseract.image_to_boxes(image_obj, lang=lang)
    with open(f'{base_output_name}_boundingBoxes.txt', 'w') as fp: 
        fp.write(text)
    # 提取详细数据
    text = pytesseract.image_to_data(image_obj, lang=lang)
    with open(f'{base_output_name}_data.txt', 'w') as fp: 
        fp.write(text)
    # 提取OSD信息
    text = pytesseract.image_to_osd(image_obj)
    with open(f'{base_output_name}_osd.txt', 'w') as fp: 
        fp.write(text)
    # 生成PDF
    pdf = pytesseract.image_to_pdf_or_hocr(image_obj, extension='pdf', lang=lang)
    with open(f'{base_output_name}.pdf', 'w+b') as f: 
        f.write(pdf)
    # 生成HOCR XML
    hocr = pytesseract.image_to_pdf_or_hocr(image_obj, extension='hocr', lang=lang)
    with open(f'{base_output_name}.hocr.xml', 'w+b') as f: 
        f.write(hocr)
    # 生成ALTO XML
    alto = pytesseract.image_to_alto_xml(image_obj)
    with open(f'{base_output_name}.alto.xml', 'w+b') as f: 
        f.write(alto)

if __name__ == '__main__':
    file_path = './radio_lomb_300.tiff' # 替换为你的文件路径,支持普通图片、单页TIFF、PDF
    file_ext = os.path.splitext(file_path)[1].lower()
    
    if file_ext == '.pdf':
        # PDF按页处理,单页PDF直接取第一页即可
        pages = convert_from_path(file_path)
        for idx, page in enumerate(pages):
            ocr_process(page, f'test_ocr_pdf_page_{idx+1}')
    else:
        # 普通图片、TIFF直接读取处理
        image = Image.open(file_path)
        ocr_process(image, 'test_ocr')

如果修复依赖后还是遇到TIFF读取问题,也可以直接把文件路径传给pytesseract的对应方法,不需要手动调用Image.open,pytesseract内部会自动适配格式读取。

内容的提问来源于stack exchange,提问作者user15235831

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 19:09:04