批量提取指定PDF页面并按位置标识合并命名的技术实现问询
批量PDF页面提取与合并方案(基于3位位置标识)
需求说明:
处理两类命名格式为Agreement1_<3位位置标识>.pdf和Agreement2_<3位位置标识>.pdf的文件,对每个位置执行:
- 从
Agreement1_*.pdf提取第3页- 从
Agreement2_*.pdf提取第1页- 合并两页生成
<位置标识>_combined.pdf
现有单文件处理命令:pdftk A=Agreement1_LOC.pdf B=Agreement2_LOC.pdf cat A3 B1 output LOC_combined.pdf,需批量处理约500个位置。
方案1:Python实现(无需依赖PDFtk)
核心思路
- 遍历目标目录筛选
Agreement1_*.pdf文件 - 从文件名提取3位位置标识
- 校验对应
Agreement2_*.pdf是否存在 - 使用
PyPDF2库完成页面提取与合并
代码实现
import os from PyPDF2 import PdfReader, PdfWriter # 替换为你的PDF文件所在目录 target_dir = "./" for filename in os.listdir(target_dir): if filename.startswith("Agreement1_") and filename.endswith(".pdf"): # 提取3位位置标识 loc_id = filename.split("_")[1].replace(".pdf", "") if len(loc_id) != 3: print(f"跳过不符合规则的文件:{filename}") continue # 拼接源文件路径 agreement1_path = os.path.join(target_dir, filename) agreement2_path = os.path.join(target_dir, f"Agreement2_{loc_id}.pdf") # 检查Agreement2文件是否存在 if not os.path.exists(agreement2_path): print(f"未找到对应文件:{agreement2_path}") continue try: reader1 = PdfReader(agreement1_path) reader2 = PdfReader(agreement2_path) writer = PdfWriter() # *注意:PyPDF2的页面索引从0开始,第3页对应索引2* if len(reader1.pages) >= 3: writer.add_page(reader1.pages[2]) else: print(f"{agreement1_path} 页数不足3页,跳过") continue if len(reader2.pages) >= 1: writer.add_page(reader2.pages[0]) else: print(f"{agreement2_path} 无有效页面,跳过") continue # 保存合并文件 output_path = os.path.join(target_dir, f"{loc_id}_combined.pdf") with open(output_path, "wb") as output_file: writer.write(output_file) print(f"已生成:{output_path}") except Exception as e: print(f"处理{loc_id}出错:{str(e)}")
使用步骤
- 安装依赖:
pip install PyPDF2 - 修改
target_dir为实际文件目录 - 运行脚本自动处理所有符合规则的文件
方案2:PowerShell实现(依赖PDFtk Pro)
核心思路
- 遍历目录下
Agreement1_*.pdf文件 - 用正则匹配提取3位位置标识
- 调用PDFtk命令执行页面提取与合并
代码实现
# 替换为你的PDF文件所在目录 $targetDir = "." Get-ChildItem -Path $targetDir -Filter "Agreement1_*.pdf" | ForEach-Object { # 正则匹配3位位置标识 if ($_.Name -match 'Agreement1_(\w{3})\.pdf') { $locId = $matches[1] $agreement1Path = $_.FullName $agreement2Path = Join-Path -Path $targetDir -ChildPath "Agreement2_$locId.pdf" if (Test-Path -Path $agreement2Path) { $outputPath = Join-Path -Path $targetDir -ChildPath "$locId_combined.pdf" # 调用PDFtk命令 pdftk A="$agreement1Path" B="$agreement2Path" cat A3 B1 output "$outputPath" if ($LASTEXITCODE -eq 0) { Write-Host "已成功生成:$outputPath" } else { Write-Host "处理$locId失败,退出码:$LASTEXITCODE" } } else { Write-Host "未找到对应文件:$agreement2Path" } } else { Write-Host "文件名不符合规则,跳过:$($_.Name)" } }
使用步骤
- 确保PDFtk Pro已加入系统环境变量(可在PowerShell中输入
pdftk验证) - 修改
$targetDir为实际文件目录 - 保存为
.ps1文件,右键选择「使用PowerShell运行」
方案3:批处理(CMD)实现(依赖PDFtk Pro)
核心思路
- 遍历目录下
Agreement1_*.pdf文件 - 通过字符串截取提取3位位置标识
- 调用PDFtk命令批量处理
代码实现
@echo off setlocal enabledelayedexpansion :: 替换为你的PDF文件所在目录 set "target_dir=." for %%f in ("%target_dir%\Agreement1_*.pdf") do ( set "filename=%%~nf" :: 截取3位位置标识(Agreement1_共11个字符,从第11位开始取3位) set "loc_id=!filename:~10,3!" if "!loc_id!"=="" ( echo 跳过不符合规则的文件:%%~nxf goto :next ) set "agreement2=%target_dir%\Agreement2_!loc_id!.pdf" if not exist "!agreement2!" ( echo 未找到对应文件:!agreement2! goto :next ) set "output=%target_dir%\!loc_id!_combined.pdf" pdftk A="%%f" B="!agreement2!" cat A3 B1 output "!output!" if errorlevel 1 ( echo 处理!loc_id!失败 ) else ( echo 已生成合并文件:!output! ) :next ) pause
使用步骤
- 确保PDFtk Pro在系统PATH中,或把
pdftk.exe放到脚本同一目录 - 修改
target_dir为实际文件目录 - 保存为
.bat文件,双击运行
内容的提问来源于stack exchange,提问作者61912
相关产品推荐
相关产品推荐

