You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

批量提取指定PDF页面并按位置标识合并命名的技术实现问询

批量PDF页面提取与合并方案(基于3位位置标识)

需求说明:
处理两类命名格式为 Agreement1_<3位位置标识>.pdf 和 Agreement2_<3位位置标识>.pdf 的文件,对每个位置执行:

  1. 从Agreement1_*.pdf提取第3页
  2. 从Agreement2_*.pdf提取第1页
  3. 合并两页生成 <位置标识>_combined.pdf
    现有单文件处理命令:pdftk A=Agreement1_LOC.pdf B=Agreement2_LOC.pdf cat A3 B1 output LOC_combined.pdf,需批量处理约500个位置。

方案1:Python实现(无需依赖PDFtk)

核心思路

  • 遍历目标目录筛选Agreement1_*.pdf文件
  • 从文件名提取3位位置标识
  • 校验对应Agreement2_*.pdf是否存在
  • 使用PyPDF2库完成页面提取与合并

代码实现

import os
from PyPDF2 import PdfReader, PdfWriter

# 替换为你的PDF文件所在目录
target_dir = "./"

for filename in os.listdir(target_dir):
    if filename.startswith("Agreement1_") and filename.endswith(".pdf"):
        # 提取3位位置标识
        loc_id = filename.split("_")[1].replace(".pdf", "")
        if len(loc_id) != 3:
            print(f"跳过不符合规则的文件:{filename}")
            continue
        
        # 拼接源文件路径
        agreement1_path = os.path.join(target_dir, filename)
        agreement2_path = os.path.join(target_dir, f"Agreement2_{loc_id}.pdf")
        
        # 检查Agreement2文件是否存在
        if not os.path.exists(agreement2_path):
            print(f"未找到对应文件:{agreement2_path}")
            continue
        
        try:
            reader1 = PdfReader(agreement1_path)
            reader2 = PdfReader(agreement2_path)
            writer = PdfWriter()
            
            # *注意:PyPDF2的页面索引从0开始,第3页对应索引2*
            if len(reader1.pages) >= 3:
                writer.add_page(reader1.pages[2])
            else:
                print(f"{agreement1_path} 页数不足3页,跳过")
                continue
            
            if len(reader2.pages) >= 1:
                writer.add_page(reader2.pages[0])
            else:
                print(f"{agreement2_path} 无有效页面,跳过")
                continue
            
            # 保存合并文件
            output_path = os.path.join(target_dir, f"{loc_id}_combined.pdf")
            with open(output_path, "wb") as output_file:
                writer.write(output_file)
            
            print(f"已生成:{output_path}")
        
        except Exception as e:
            print(f"处理{loc_id}出错:{str(e)}")

使用步骤

  1. 安装依赖:pip install PyPDF2
  2. 修改target_dir为实际文件目录
  3. 运行脚本自动处理所有符合规则的文件

方案2:PowerShell实现(依赖PDFtk Pro)

核心思路

  • 遍历目录下Agreement1_*.pdf文件
  • 用正则匹配提取3位位置标识
  • 调用PDFtk命令执行页面提取与合并

代码实现

# 替换为你的PDF文件所在目录
$targetDir = "."

Get-ChildItem -Path $targetDir -Filter "Agreement1_*.pdf" | ForEach-Object {
    # 正则匹配3位位置标识
    if ($_.Name -match 'Agreement1_(\w{3})\.pdf') {
        $locId = $matches[1]
        $agreement1Path = $_.FullName
        $agreement2Path = Join-Path -Path $targetDir -ChildPath "Agreement2_$locId.pdf"
        
        if (Test-Path -Path $agreement2Path) {
            $outputPath = Join-Path -Path $targetDir -ChildPath "$locId_combined.pdf"
            
            # 调用PDFtk命令
            pdftk A="$agreement1Path" B="$agreement2Path" cat A3 B1 output "$outputPath"
            
            if ($LASTEXITCODE -eq 0) {
                Write-Host "已成功生成:$outputPath"
            } else {
                Write-Host "处理$locId失败,退出码:$LASTEXITCODE"
            }
        } else {
            Write-Host "未找到对应文件:$agreement2Path"
        }
    } else {
        Write-Host "文件名不符合规则,跳过:$($_.Name)"
    }
}

使用步骤

  1. 确保PDFtk Pro已加入系统环境变量(可在PowerShell中输入pdftk验证)
  2. 修改$targetDir为实际文件目录
  3. 保存为.ps1文件,右键选择「使用PowerShell运行」

方案3:批处理(CMD)实现(依赖PDFtk Pro)

核心思路

  • 遍历目录下Agreement1_*.pdf文件
  • 通过字符串截取提取3位位置标识
  • 调用PDFtk命令批量处理

代码实现

@echo off
setlocal enabledelayedexpansion

:: 替换为你的PDF文件所在目录
set "target_dir=."

for %%f in ("%target_dir%\Agreement1_*.pdf") do (
    set "filename=%%~nf"
    :: 截取3位位置标识(Agreement1_共11个字符,从第11位开始取3位)
    set "loc_id=!filename:~10,3!"
    
    if "!loc_id!"=="" (
        echo 跳过不符合规则的文件:%%~nxf
        goto :next
    )
    
    set "agreement2=%target_dir%\Agreement2_!loc_id!.pdf"
    if not exist "!agreement2!" (
        echo 未找到对应文件:!agreement2!
        goto :next
    )
    
    set "output=%target_dir%\!loc_id!_combined.pdf"
    
    pdftk A="%%f" B="!agreement2!" cat A3 B1 output "!output!"
    
    if errorlevel 1 (
        echo 处理!loc_id!失败
    ) else (
        echo 已生成合并文件:!output!
    )
    
    :next
)

pause

使用步骤

  1. 确保PDFtk Pro在系统PATH中,或把pdftk.exe放到脚本同一目录
  2. 修改target_dir为实际文件目录
  3. 保存为.bat文件,双击运行

内容的提问来源于stack exchange,提问作者61912

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 03:10:16