You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求提取PDF方括号内文本的JavaScript/PowerShell脚本帮助

解决方案:提取PDF中方括号内的文档ID

一、修正后的Adobe Acrobat JavaScript脚本

原脚本仅匹配整单词包裹的方括号内容,且文本提取逻辑存在局限,无法覆盖方括号跨单词或整段文本中的匹配场景。以下是优化后的版本,通过正则表达式批量提取所有符合格式的文档ID:

// 替换为你的PDF文件绝对路径(Windows路径用双反斜杠\\)
var filePath = "C:\\Users\\xxx\\Documents\\target.pdf";
// 替换为输出文件路径(支持.txt或.csv格式)
var outputPath = "C:\\Users\\xxx\\Documents\\doc_ids.txt";

try {
    var doc = app.open(filePath);
    if (!doc) {
        app.alert("无法打开PDF文件,请检查路径是否正确。");
        throw new Error("PDF文件打开失败");
    }

    var results = [];
    // 正则表达式匹配所有方括号内的内容,支持跨单词场景
    var idRegex = /\[([^\]]+)\]/g;

    for (var i = 0; i < doc.numPages; i++) {
        // 获取当前页完整文本
        var pageText = doc.getPageText(i);
        var matchResult;
        // 遍历当前页所有匹配项
        while ((matchResult = idRegex.exec(pageText)) !== null) {
            results.push(matchResult[1]);
        }
    }

    // 写入结果文件
    var outputFile = new File(outputPath);
    outputFile.open("w");
    // 若需CSV格式,将换行符改为逗号:results.join(",")
    outputFile.write(results.join("\n"));
    outputFile.close();

    doc.close();
    app.alert("提取完成,结果已保存至:" + outputPath);
} catch (e) {
    app.alert("执行出错:" + e.message);
    if (typeof doc !== "undefined" && doc) {
        doc.close();
    }
}

使用步骤:

  1. 打开Adobe Acrobat Pro/DC并加载目标PDF
  2. 按Ctrl+J调出JavaScript控制台
  3. 粘贴上述代码,修改filePath和outputPath为实际路径
  4. 点击控制台的运行按钮(▶️)执行脚本

二、PowerShell脚本(依赖Adobe Acrobat,无需额外安装软件)

若习惯使用PowerShell,可通过调用Adobe Acrobat的COM组件完成提取:

# 替换为你的PDF文件路径
$pdfPath = "C:\Users\xxx\Documents\target.pdf"
# 替换为输出文件路径
$outputPath = "C:\Users\xxx\Documents\doc_ids.txt"

# 初始化Acrobat COM对象
$acroApp = New-Object -ComObject AcroExch.App
$acroDoc = New-Object -ComObject AcroExch.PDDoc

try {
    if (-not $acroDoc.Open($pdfPath)) {
        Write-Error "无法打开PDF文件,请检查路径是否正确。"
        exit 1
    }

    $results = @()
    $idRegex = [regex]::new('\[([^\]]+)\]')

    # 遍历所有PDF页面
    for ($pageIndex = 0; $pageIndex -lt $acroDoc.GetNumPages(); $pageIndex++) {
        $page = $acroDoc.AcquirePage($pageIndex)
        $pageText = $page.GetText()
        $matches = $idRegex.Matches($pageText)
        
        foreach ($match in $matches) {
            $results += $match.Groups[1].Value
        }

        $acroDoc.ReleasePage($pageIndex)
    }

    # 写入结果到文件
    $results | Out-File -FilePath $outputPath -Encoding utf8

    Write-Host "提取完成,结果已保存至:$outputPath"
} catch {
    Write-Error "执行出错:$_"
} finally {
    # 清理COM对象,避免内存泄漏
    if ($acroDoc.IsOpen()) {
        $acroDoc.Close()
    }
    $acroApp.Exit()
    [System.Runtime.Interopservices.Marshal]::ReleaseComObject($acroDoc) | Out-Null
    [System.Runtime.Interopservices.Marshal]::ReleaseComObject($acroApp) | Out-Null
    [System.GC]::Collect()
    [System.GC]::WaitForPendingFinalizers()
}

使用步骤:

  1. 打开PowerShell(确保有权限访问目标文件路径)
  2. 修改脚本中的$pdfPath和$outputPath为实际路径
  3. 粘贴脚本并回车执行

内容的提问来源于stack exchange,提问作者Leroy P

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 11:15:44