You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python pdf2image合并多帧TIFF时遇类型错误的解决咨询

问题描述

使用pdf2image将PDF导出为多页TIFF的基础代码运行正常:

images = convert_from_path(p_filePath, dpi = DPI, poppler_path = POPPLER_PATH)
images[0].save(os.path.join(p_targetPath, outputFilename), self.realType, save_all=True, 
                        append_images=images[1:], compression='tiff_adobe_deflate')

但尝试将新PDF生成的图像合并到已有同名TIFF时,两种实现方式均抛出异常:

TypeError: int() argument must be a string, a bytes-like object or a real number, not 'NoneType'

异常根源为TiffImagePlugin.py中尝试获取IMAGEWIDTH标签时得到None值。

已确认原有TIFF文件无损坏,用seek统计页数正常,且用tifftools合并TIFF也能成功。观察到合并时的图像列表包含两种不同类型:PIL.TiffImagePlugin.TiffImageFile(来自已有TIFF)和PIL.PpmImagePlugin.PpmImageFile(来自convert_from_path输出)。

原因分析

问题出在混合使用不同PIL图像子类进行多页TIFF保存:
当以TiffImageFile作为主图像调用save_all时,Pillow会启用TIFF专属的处理逻辑遍历所有图像。而PpmImageFile对象没有TIFF格式特有的IMAGEWIDTH元数据标签,导致在执行seek操作时,代码尝试将None转换为整数,触发TypeError。

修复方案

核心是将所有图像统一转换为普通的RGB模式Image对象,避免混合不同子类。以下是两种实现的修改版本:

方案1:修改mergeTiff函数

from PIL import Image, ImageSequence
import os
from pdf2image import convert_from_path

def mergeTiff(self, p_imageList, p_destDir, p_outputFilename):
    imageList = []
    tiff_path = os.path.join(p_destDir, p_outputFilename)
    if os.path.isfile(tiff_path):
        tiff_img = Image.open(tiff_path)
        for page in ImageSequence.Iterator(tiff_img):
            # 将TIFF页面转换为RGB模式的普通图像
            imageList.append(page.convert("RGB"))
        tiff_img.close()
    # 将新生成的Ppm图像也转换为RGB模式
    converted_new_images = [img.convert("RGB") for img in p_imageList]
    imageList.extend(converted_new_images)
    return imageList

# 调用示例
images = convert_from_path(p_filePath, dpi=DPI, poppler_path=POPPLER_PATH)
images = self.mergeTiff(images, p_targetPath, outputFilename)
if images:
    images[0].save(os.path.join(p_targetPath, outputFilename), 
                   self.realType, 
                   save_all=True, 
                   append_images=images[1:], 
                   compression='tiff_adobe_deflate')

方案2:直接合并转换后的图像列表

from PIL import Image, ImageSequence
import os
from pdf2image import convert_from_path

tiff_path = os.path.join(p_destDir, p_outputFilename)
existing_images = []

# 读取已有TIFF并转换所有页面为RGB
if os.path.isfile(tiff_path):
    tiff_img = Image.open(tiff_path)
    for page in ImageSequence.Iterator(tiff_img):
        existing_images.append(page.convert("RGB"))
    tiff_img.close()

# 获取新PDF的图像并转换为RGB
new_images = [img.convert("RGB") for img in convert_from_path(p_filePath, dpi=DPI, poppler_path=POPPLER_PATH)]

# 合并并保存
all_images = existing_images + new_images
if all_images:
    all_images[0].save(os.path.join(p_targetPath, outputFilename), 
                       self.realType, 
                       save_all=True, 
                       append_images=all_images[1:], 
                       compression='tiff_adobe_deflate')

关键说明

  • 转换为RGB模式:统一图像格式,避免不同模式(如灰度L模式)的兼容问题;同时将子类对象转换为通用的Image对象,让Pillow使用通用图像处理逻辑保存多页TIFF,不会触发TIFF专属的元数据读取逻辑。
  • 关闭TIFF文件:显式关闭打开的TIFF文件,避免资源泄漏。

内容的提问来源于stack exchange,提问作者Joseph Orfeuil

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 04:04:56