如何捕获PIL库Image.save函数输出的通用错误文本?
问题背景
使用PIL将TIFF文件批量转换为PDF时,遇到损坏图像会触发底层库输出类似以下的错误信息,但无法通过Python的warnings模块或普通stdout/stderr重定向捕获这些信息,从而难以定位问题图像:
Fax4Decode: Uncompressed data (not supported) at line 1333 of strip 0 (x 770). Fax4Decode: Uncompressed data (not supported) at line 1334 of strip 0 (x 0). Fax4Decode: Bad code word at line 1613 of strip 0 (x 1694).
错误产生于以下代码行:
images[0].save(pf_path, "PDF", resolution=100.0, save_all=True, append_images=images[1:])
这些信息来自底层libtiff的C级输出,Python的常规捕获方式无法覆盖,以下是三种可行的解决方案:
解决方案1:通过ctypes捕获libtiff错误输出
libtiff提供了错误回调接口,可通过ctypes覆盖默认处理函数,将C层错误信息转存到Python变量中:
import ctypes from PIL import Image # 加载libtiff库,路径需根据系统调整: # Linux: "libtiff.so" | macOS: "/usr/lib/libtiff.dylib" | Windows: "tiff.dll" libtiff = ctypes.CDLL("libtiff.so") # 定义错误回调函数类型 ErrorHandler = ctypes.CFUNCTYPE(None, ctypes.c_char_p, ctypes.c_int, ctypes.c_char_p, ctypes.c_va_list) # 存储错误信息的容器 tiff_errors = [] def tiff_error_handler(module, level, fmt, ap): buf = ctypes.create_string_buffer(1024) ctypes.vsprintf(buf, fmt, ap) error_msg = buf.value.decode('utf-8') tiff_errors.append(f"[{module.decode('utf-8')}] {error_msg}") # 注册回调函数 c_error_handler = ErrorHandler(tiff_error_handler) libtiff.TIFFSetErrorHandler(c_error_handler) libtiff.TIFFSetWarningHandler(c_error_handler) # 执行转换逻辑 target_files = ["tiff_file1.tiff", "corrupted_tiff.tiff"] images = [Image.open(f) for f in target_files] output_pdf = "result.pdf" tiff_errors.clear() try: images[0].save(output_pdf, "PDF", resolution=100.0, save_all=True, append_images=images[1:]) except Exception as e: print(f"转换失败: {str(e)}") # 输出捕获的错误,关联问题文件 if tiff_errors: print("捕获到TIFF处理错误:") for err in tiff_errors: print(err) # 此处可结合target_files列表定位具体问题文件
解决方案2:用subprocess隔离进程捕获stderr
将单文件转换逻辑放到独立脚本中,通过subprocess调用并捕获子进程的stderr,确保完整获取底层C输出:
子脚本 tiff_convert_worker.py
from PIL import Image import sys def main(): tiff_path = sys.argv[1] pdf_path = sys.argv[2] try: img = Image.open(tiff_path) img.save(pdf_path, "PDF", resolution=100.0) except Exception as e: print(f"转换异常: {str(e)}", file=sys.stderr) sys.exit(1) if __name__ == "__main__": main()
主调用脚本
import subprocess import os tiff_files = ["file1.tiff", "corrupted_file.tiff", "file2.tiff"] output_dir = "pdf_out" os.makedirs(output_dir, exist_ok=True) for tiff_path in tiff_files: pdf_name = f"{os.path.splitext(os.path.basename(tiff_path))[0]}.pdf" pdf_path = os.path.join(output_dir, pdf_name) proc_result = subprocess.run( ["python", "tiff_convert_worker.py", tiff_path, pdf_path], capture_output=True, text=True ) if proc_result.stderr or proc_result.returncode != 0: print(f"处理文件 {tiff_path} 出现问题:") print(proc_result.stderr)
解决方案3:提前用Image.verify()检查图像
在转换前对每个TIFF文件调用verify(),提前过滤明显损坏的图像:
from PIL import Image tiff_files = ["file1.tiff", "corrupted_file.tiff", "file2.tiff"] problem_files = [] valid_images = [] for tiff_path in tiff_files: try: with Image.open(tiff_path) as img: img.verify() # 检查文件完整性 valid_images.append(Image.open(tiff_path)) # 重新打开(verify会重置文件指针) except Exception as e: problem_files.append((tiff_path, str(e))) print(f"文件 {tiff_path} 验证失败:{e}") # 对有效图像执行批量转换 if valid_images: valid_images[0].save("final.pdf", "PDF", resolution=100.0, save_all=True, append_images=valid_images[1:])
内容的提问来源于stack exchange,提问作者Steven M.
相关产品推荐
相关产品推荐

