使用Python转换超宽扫描PDF为图片时仅生成1像素白块的问题
解决超宽PDF转图片仅生成1像素宽白块的问题
问题根源
你的问题出在pdf2image依赖的poppler工具链上,它对超宽页面(如30000px宽度)有默认尺寸限制,自动截断后生成了异常的1像素宽图像,问题并非出在PIL或后续保存步骤。
改进方案
通过给convert_from_path传递poppler_options参数,强制poppler按原页面尺寸比例渲染,避免自动截断:
- 指定输出宽度并保持高度自适应比例
- 禁用cropbox,确保完整渲染原始页面范围
修改后的完整代码
from PIL import Image from pdf2image import pdfinfo_from_path, convert_from_path from pathlib import Path # 解除PIL的像素数量限制 Image.MAX_IMAGE_PIXELS = None def pdf_to_img(pdf_path, dpi=500): """ Convert the PDF file to JPEG images """ pdf_name = Path(pdf_path).stem # 用Path工具更健壮地提取文件名 pdf_info = pdfinfo_from_path(pdf_path, userpw=None, poppler_path=None) page_nb = pdf_info["Pages"] step = 2 try: for img_nb in range(1, page_nb + 1, step): batch_pages = convert_from_path( pdf_path, dpi=dpi, first_page=img_nb, last_page=min(img_nb + step - 1, page_nb), # 关键参数:强制按原宽度渲染,高度自适应,禁用裁剪框 poppler_options=[ "-scale-to-x", "30000", # 匹配你的PDF页面宽度 "-scale-to-y", "-1", # 高度按比例自动计算 "-use-cropbox", "false" # 使用原始页面范围而非裁剪区域 ] ) # 用enumerate遍历,避免手动自增页码的逻辑混乱 for idx, page in enumerate(batch_pages): current_page_num = img_nb + idx save_img(page, f"{pdf_name}_{current_page_num:04d}.jpg") except Exception as e: print(f"[pdf_to_img] Failed to convert {pdf_name}.pdf to images:\n{e} ({e.__class__.__name__})") def save_img( img, img_filename, img_path=Path("./output"), error_msg="Failed to save img", max_dim=2500, img_format="JPEG", ): try: # 确保输出目录存在,避免保存失败 img_path.mkdir(exist_ok=True) if img.width > max_dim or img.height > max_dim: img.thumbnail( (max_dim, max_dim), Image.Resampling.LANCZOS ) img.save(img_path / img_filename, format=img_format) return True except Exception as e: print(f"[save_img] {error_msg}:\n{e} ({e.__class__.__name__})") return False
额外优化说明
- 用
Path(pdf_path).stem替代字符串分割,适配不同操作系统的路径格式 - 用
enumerate批量处理页面,避免手动自增页码引发的逻辑错误 - 添加
img_path.mkdir(exist_ok=True),自动创建输出目录,减少异常场景
内容的提问来源于stack exchange,提问作者Seglinglin
相关产品推荐
相关产品推荐

