You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyPDF2 1.26.0合并PDF页面后出现空白页问题求助

交替合并两个PDF后出现空白页的修正方案

需求:将两个PDF文件按页交替合并(第一个PDF第1页→第二个PDF第1页→第一个PDF第2页→第二个PDF第2页,依此类推),使用公司限定的PyPDF2 1.26.0版本,基于Alteryx编写Python脚本,但合并后的PDF全为空白页,现有脚本如下:

from ayx import Alteryx
from PyPDF2 import PdfFileReader, PdfFileWriter
import os
import PyPDF2


directory_path =  Alteryx.read('#1').iloc[0,0] #this projected is done in Alteryx, my path is to folder #with two PDF files
pdf_files = [file for file in os.listdir(directory_path)]
output_pdf = PdfFileWriter()



for i in range(0, min(len(pdf_files), len(pdf_files)-  len(pdf_files)%  2), 2):
    with open(os.path.join(directory_path,pdf_files[i]), 'rb') as file1,open(os.path.join(directory_path, pdf_files[i+1]),'rb') as file2:
        reader1 = PdfFileReader(file1)
        reader2 = PdfFileReader(file2)
        
        for page_num in range(max(reader1.getNumPages(), reader2.getNumPages())):
            if page_num < reader1.getNumPages():
                output_pdf.addPage(reader1.getPage(page_num))
            if page_num < reader2.getNumPages():
                output_pdf.addPage(reader2.getPage(page_num))
                
output_file_path = os.path.join(directory_path, 'merged.pdf')
with open(output_file_path, 'wb') as output_file:
    output_pdf.write(output_file)

错误原因分析

  • 文件列表不可控:os.listdir()返回的文件顺序由系统决定,可能混入非PDF文件,或两个目标PDF的顺序不符合预期,导致读取错误文件。
  • PyPDF2版本兼容问题:1.26.0版本的PdfFileReader对部分PDF的读取需要关闭严格模式,否则可能无法正常解析页面内容。
  • 循环逻辑冗余:外层循环针对多组PDF设计,但实际只需要处理两个文件,冗余逻辑可能引发意外问题。
  • 缺乏错误校验:未校验读取的文件是否为有效PDF,若误读非PDF文件会导致写入空白页面。

修正后的脚本

from ayx import Alteryx
from PyPDF2 import PdfFileReader, PdfFileWriter
import os

# 读取Alteryx传入的文件夹路径
directory_path = Alteryx.read('#1').iloc[0, 0]

# 过滤文件夹中的PDF文件,确保只处理目标文件
pdf_files = [
    file for file in os.listdir(directory_path)
    if file.lower().endswith('.pdf')
]

# 校验是否有且仅有两个PDF文件
if len(pdf_files) != 2:
    raise ValueError("目标文件夹中必须且只能有两个PDF文件")

# 可根据需求调整文件顺序,比如按文件名排序
pdf_files.sort()

output_pdf = PdfFileWriter()

# 打开两个PDF文件,添加strict=False兼容更多PDF格式
with open(os.path.join(directory_path, pdf_files[0]), 'rb') as file1, \
     open(os.path.join(directory_path, pdf_files[1]), 'rb') as file2:
    
    reader1 = PdfFileReader(file1, strict=False)
    reader2 = PdfFileReader(file2, strict=False)
    
    # 获取两个PDF的最大页数,确保遍历完所有页面
    max_pages = max(reader1.getNumPages(), reader2.getNumPages())
    
    for page_num in range(max_pages):
        # 交替添加页面:先加第一个PDF的当前页(如果存在)
        if page_num < reader1.getNumPages():
            output_pdf.addPage(reader1.getPage(page_num))
        # 再加第二个PDF的当前页(如果存在)
        if page_num < reader2.getNumPages():
            output_pdf.addPage(reader2.getPage(page_num))

# 输出合并后的PDF
output_file_path = os.path.join(directory_path, 'merged.pdf')
with open(output_file_path, 'wb') as output_file:
    output_pdf.write(output_file)

关键修改说明

  • 过滤PDF文件:只保留后缀为.pdf的文件,避免误读非PDF内容。
  • 文件顺序校验与排序:确保文件夹内只有两个PDF,并通过排序固定合并顺序(可根据实际需求调整排序规则)。
  • 关闭严格模式:给PdfFileReader添加strict=False参数,兼容1.26.0版本中对部分PDF的解析问题。
  • 简化逻辑:移除冗余的外层循环,直接针对两个PDF处理,逻辑更清晰。
  • 增加校验:提前检查文件数量,避免后续无意义的执行。

内容的提问来源于stack exchange,提问作者MaksymMD

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 15:44:55