You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的PdfFileMerger按前缀合并PDF时重复页面问题排查

问题描述

有一个存放单页PDF的目录,文件命名规则为「前缀+数字」(如A_001.pdf、B_004.pdf),需要将同前缀的PDF合并为对应多页PDF(如A.pdf包含所有A前缀的单页PDF,B.pdf同理)。当前Python脚本虽能生成目标文件,但所有输出文件都重复包含了所有类别的PDF页面,无法正确合并。

原脚本代码

import PyPDF2, os
from PyPDF2 import PdfFileReader, PdfFileWriter, PdfFileMerger
from pathlib import Path

single_file_dir = r'Y:\Python\Single_PDFs'
binder_file_dir = r'Y:\Python\Combined_PDFs'

# get list of all files in the single PDF directory
single_file_list = []
for file in os.listdir(single_file_dir):
    if file.endswith(".pdf"):
        single_file_list.append(single_file_dir + "\\" + file)

print(single_file_list)

# get the file names for the output multi page pdfs

file_name_list = []
for file in single_file_list:
    name = os.path.basename(file)
    new_name = name[:-8]
    file_name_list.append(new_name)
    unique_file_name_list = list(set(file_name_list))

merger = PdfFileMerger()

print(unique_file_name_list)

#try to match input single file name to output file name
for file in single_file_list:
    for name in unique_file_name_list:
        if name in file:
            merger.append(file)
            merger.write(binder_file_dir + "\\" + name + ".pdf")

问题原因

  1. 复用同一个合并实例:全程只用了一个PdfFileMerger对象,每次添加文件后直接写入,导致后续写入的文件会包含之前所有添加过的页面,最终所有输出文件都累积了全部页面。
  2. 循环逻辑错误:嵌套循环会让每个文件匹配到所有符合的前缀(实际每个文件只属于一个前缀),重复执行append和write操作,加剧了页面重复的问题。
  3. 未处理文件排序:os.listdir返回的文件顺序不固定,可能导致合并后的PDF页面顺序混乱。

修正后的代码

import os
from PyPDF2 import PdfFileMerger

single_file_dir = r'Y:\Python\Single_PDFs'
binder_file_dir = r'Y:\Python\Combined_PDFs'

# 确保输出目录存在
os.makedirs(binder_file_dir, exist_ok=True)

# 按前缀分组文件
pdf_groups = {}
for file in os.listdir(single_file_dir):
    if file.endswith(".pdf"):
        # 提取前缀(假设分隔符是下划线)
        prefix = file.split("_")[0]
        file_path = os.path.join(single_file_dir, file)
        # 将文件加入对应前缀的组
        if prefix not in pdf_groups:
            pdf_groups[prefix] = []
        pdf_groups[prefix].append(file_path)

# 遍历每个分组,合并PDF
for prefix, files in pdf_groups.items():
    # 对每个分组单独创建合并实例
    merger = PdfFileMerger()
    # 按文件名排序,确保页面顺序正确(比如A_001、A_002...)
    sorted_files = sorted(files, key=lambda x: os.path.basename(x))
    for file in sorted_files:
        merger.append(file)
    # 写入合并后的文件
    output_path = os.path.join(binder_file_dir, f"{prefix}.pdf")
    merger.write(output_path)
    merger.close()  # 关闭合并器释放资源

print("PDF合并完成")

关键改进点

  • 按前缀分组:先将所有文件按前缀归类,避免交叉合并。
  • 独立合并实例:每个前缀分组使用单独的PdfFileMerger,确保各合并任务互不干扰。
  • 文件排序:对每个分组内的文件按名称排序,保证合并后的页面顺序符合数字递增的要求。
  • 目录检查:添加os.makedirs确保输出目录存在,避免写入失败。

内容的提问来源于stack exchange,提问作者Cor_Novum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.17 19:53:16