使用Python的PdfFileMerger按前缀合并PDF时重复页面问题排查
问题描述
有一个存放单页PDF的目录,文件命名规则为「前缀+数字」(如A_001.pdf、B_004.pdf),需要将同前缀的PDF合并为对应多页PDF(如A.pdf包含所有A前缀的单页PDF,B.pdf同理)。当前Python脚本虽能生成目标文件,但所有输出文件都重复包含了所有类别的PDF页面,无法正确合并。
原脚本代码
import PyPDF2, os from PyPDF2 import PdfFileReader, PdfFileWriter, PdfFileMerger from pathlib import Path single_file_dir = r'Y:\Python\Single_PDFs' binder_file_dir = r'Y:\Python\Combined_PDFs' # get list of all files in the single PDF directory single_file_list = [] for file in os.listdir(single_file_dir): if file.endswith(".pdf"): single_file_list.append(single_file_dir + "\\" + file) print(single_file_list) # get the file names for the output multi page pdfs file_name_list = [] for file in single_file_list: name = os.path.basename(file) new_name = name[:-8] file_name_list.append(new_name) unique_file_name_list = list(set(file_name_list)) merger = PdfFileMerger() print(unique_file_name_list) #try to match input single file name to output file name for file in single_file_list: for name in unique_file_name_list: if name in file: merger.append(file) merger.write(binder_file_dir + "\\" + name + ".pdf")
问题原因
- 复用同一个合并实例:全程只用了一个
PdfFileMerger对象,每次添加文件后直接写入,导致后续写入的文件会包含之前所有添加过的页面,最终所有输出文件都累积了全部页面。 - 循环逻辑错误:嵌套循环会让每个文件匹配到所有符合的前缀(实际每个文件只属于一个前缀),重复执行append和write操作,加剧了页面重复的问题。
- 未处理文件排序:
os.listdir返回的文件顺序不固定,可能导致合并后的PDF页面顺序混乱。
修正后的代码
import os from PyPDF2 import PdfFileMerger single_file_dir = r'Y:\Python\Single_PDFs' binder_file_dir = r'Y:\Python\Combined_PDFs' # 确保输出目录存在 os.makedirs(binder_file_dir, exist_ok=True) # 按前缀分组文件 pdf_groups = {} for file in os.listdir(single_file_dir): if file.endswith(".pdf"): # 提取前缀(假设分隔符是下划线) prefix = file.split("_")[0] file_path = os.path.join(single_file_dir, file) # 将文件加入对应前缀的组 if prefix not in pdf_groups: pdf_groups[prefix] = [] pdf_groups[prefix].append(file_path) # 遍历每个分组,合并PDF for prefix, files in pdf_groups.items(): # 对每个分组单独创建合并实例 merger = PdfFileMerger() # 按文件名排序,确保页面顺序正确(比如A_001、A_002...) sorted_files = sorted(files, key=lambda x: os.path.basename(x)) for file in sorted_files: merger.append(file) # 写入合并后的文件 output_path = os.path.join(binder_file_dir, f"{prefix}.pdf") merger.write(output_path) merger.close() # 关闭合并器释放资源 print("PDF合并完成")
关键改进点
- 按前缀分组:先将所有文件按前缀归类,避免交叉合并。
- 独立合并实例:每个前缀分组使用单独的
PdfFileMerger,确保各合并任务互不干扰。 - 文件排序:对每个分组内的文件按名称排序,保证合并后的页面顺序符合数字递增的要求。
- 目录检查:添加
os.makedirs确保输出目录存在,避免写入失败。
内容的提问来源于stack exchange,提问作者Cor_Novum
相关产品推荐
相关产品推荐

