You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何捕获PIL库Image.save函数输出的通用错误文本?

问题背景

使用PIL将TIFF文件批量转换为PDF时,遇到损坏图像会触发底层库输出类似以下的错误信息,但无法通过Python的warnings模块或普通stdout/stderr重定向捕获这些信息,从而难以定位问题图像:

Fax4Decode: Uncompressed data (not supported) at line 1333 of strip 0 (x 770).
Fax4Decode: Uncompressed data (not supported) at line 1334 of strip 0 (x 0).
Fax4Decode: Bad code word at line 1613 of strip 0 (x 1694).

错误产生于以下代码行:

images[0].save(pf_path, "PDF", resolution=100.0, save_all=True, append_images=images[1:])

这些信息来自底层libtiff的C级输出,Python的常规捕获方式无法覆盖,以下是三种可行的解决方案:


解决方案1:通过ctypes捕获libtiff错误输出

libtiff提供了错误回调接口,可通过ctypes覆盖默认处理函数,将C层错误信息转存到Python变量中:

import ctypes
from PIL import Image

# 加载libtiff库,路径需根据系统调整:
# Linux: "libtiff.so" | macOS: "/usr/lib/libtiff.dylib" | Windows: "tiff.dll"
libtiff = ctypes.CDLL("libtiff.so")

# 定义错误回调函数类型
ErrorHandler = ctypes.CFUNCTYPE(None, ctypes.c_char_p, ctypes.c_int, ctypes.c_char_p, ctypes.c_va_list)

# 存储错误信息的容器
tiff_errors = []

def tiff_error_handler(module, level, fmt, ap):
    buf = ctypes.create_string_buffer(1024)
    ctypes.vsprintf(buf, fmt, ap)
    error_msg = buf.value.decode('utf-8')
    tiff_errors.append(f"[{module.decode('utf-8')}] {error_msg}")

# 注册回调函数
c_error_handler = ErrorHandler(tiff_error_handler)
libtiff.TIFFSetErrorHandler(c_error_handler)
libtiff.TIFFSetWarningHandler(c_error_handler)

# 执行转换逻辑
target_files = ["tiff_file1.tiff", "corrupted_tiff.tiff"]
images = [Image.open(f) for f in target_files]
output_pdf = "result.pdf"

tiff_errors.clear()
try:
    images[0].save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=images[1:])
except Exception as e:
    print(f"转换失败: {str(e)}")

# 输出捕获的错误,关联问题文件
if tiff_errors:
    print("捕获到TIFF处理错误:")
    for err in tiff_errors:
        print(err)
    # 此处可结合target_files列表定位具体问题文件

解决方案2:用subprocess隔离进程捕获stderr

将单文件转换逻辑放到独立脚本中,通过subprocess调用并捕获子进程的stderr,确保完整获取底层C输出:

子脚本 tiff_convert_worker.py

from PIL import Image
import sys

def main():
    tiff_path = sys.argv[1]
    pdf_path = sys.argv[2]
    try:
        img = Image.open(tiff_path)
        img.save(pdf_path, "PDF", resolution=100.0)
    except Exception as e:
        print(f"转换异常: {str(e)}", file=sys.stderr)
        sys.exit(1)

if __name__ == "__main__":
    main()

主调用脚本

import subprocess
import os

tiff_files = ["file1.tiff", "corrupted_file.tiff", "file2.tiff"]
output_dir = "pdf_out"
os.makedirs(output_dir, exist_ok=True)

for tiff_path in tiff_files:
    pdf_name = f"{os.path.splitext(os.path.basename(tiff_path))[0]}.pdf"
    pdf_path = os.path.join(output_dir, pdf_name)
    
    proc_result = subprocess.run(
        ["python", "tiff_convert_worker.py", tiff_path, pdf_path],
        capture_output=True,
        text=True
    )
    
    if proc_result.stderr or proc_result.returncode != 0:
        print(f"处理文件 {tiff_path} 出现问题:")
        print(proc_result.stderr)

解决方案3:提前用Image.verify()检查图像

在转换前对每个TIFF文件调用verify(),提前过滤明显损坏的图像:

from PIL import Image

tiff_files = ["file1.tiff", "corrupted_file.tiff", "file2.tiff"]
problem_files = []
valid_images = []

for tiff_path in tiff_files:
    try:
        with Image.open(tiff_path) as img:
            img.verify()  # 检查文件完整性
            valid_images.append(Image.open(tiff_path))  # 重新打开(verify会重置文件指针)
    except Exception as e:
        problem_files.append((tiff_path, str(e)))
        print(f"文件 {tiff_path} 验证失败:{e}")

# 对有效图像执行批量转换
if valid_images:
    valid_images[0].save("final.pdf", "PDF", resolution=100.0, save_all=True, append_images=valid_images[1:])

内容的提问来源于stack exchange,提问作者Steven M.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.25 00:42:42