You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python转换超宽扫描PDF为图片时仅生成1像素白块的问题

解决超宽PDF转图片仅生成1像素宽白块的问题

问题根源

你的问题出在pdf2image依赖的poppler工具链上,它对超宽页面(如30000px宽度)有默认尺寸限制,自动截断后生成了异常的1像素宽图像,问题并非出在PIL或后续保存步骤。

改进方案

通过给convert_from_path传递poppler_options参数,强制poppler按原页面尺寸比例渲染,避免自动截断:

  • 指定输出宽度并保持高度自适应比例
  • 禁用cropbox,确保完整渲染原始页面范围

修改后的完整代码

from PIL import Image
from pdf2image import pdfinfo_from_path, convert_from_path
from pathlib import Path

# 解除PIL的像素数量限制
Image.MAX_IMAGE_PIXELS = None

def pdf_to_img(pdf_path, dpi=500):
    """
    Convert the PDF file to JPEG images
    """
    pdf_name = Path(pdf_path).stem  # 用Path工具更健壮地提取文件名
    pdf_info = pdfinfo_from_path(pdf_path, userpw=None, poppler_path=None)
    page_nb = pdf_info["Pages"]
    step = 2
    try:
        for img_nb in range(1, page_nb + 1, step):
            batch_pages = convert_from_path(
                pdf_path,
                dpi=dpi,
                first_page=img_nb,
                last_page=min(img_nb + step - 1, page_nb),
                # 关键参数:强制按原宽度渲染,高度自适应,禁用裁剪框
                poppler_options=[
                    "-scale-to-x", "30000",  # 匹配你的PDF页面宽度
                    "-scale-to-y", "-1",     # 高度按比例自动计算
                    "-use-cropbox", "false"  # 使用原始页面范围而非裁剪区域
                ]
            )
            # 用enumerate遍历,避免手动自增页码的逻辑混乱
            for idx, page in enumerate(batch_pages):
                current_page_num = img_nb + idx
                save_img(page, f"{pdf_name}_{current_page_num:04d}.jpg")
    except Exception as e:
        print(f"[pdf_to_img] Failed to convert {pdf_name}.pdf to images:\n{e} ({e.__class__.__name__})")


def save_img(
    img,
    img_filename,
    img_path=Path("./output"),
    error_msg="Failed to save img",
    max_dim=2500,
    img_format="JPEG",
):
    try:
        # 确保输出目录存在,避免保存失败
        img_path.mkdir(exist_ok=True)
        if img.width > max_dim or img.height > max_dim:
            img.thumbnail(
                (max_dim, max_dim), Image.Resampling.LANCZOS
            )
        img.save(img_path / img_filename, format=img_format)
        return True
    except Exception as e:
        print(f"[save_img] {error_msg}:\n{e} ({e.__class__.__name__})")
    return False

额外优化说明

  • 用Path(pdf_path).stem替代字符串分割,适配不同操作系统的路径格式
  • 用enumerate批量处理页面,避免手动自增页码引发的逻辑错误
  • 添加img_path.mkdir(exist_ok=True),自动创建输出目录,减少异常场景

内容的提问来源于stack exchange,提问作者Seglinglin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 06:15:38