You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:用VBS/PowerShell/CMD批量匹配PDF内容并复制文件

PowerShell 解决方案

这个方案利用Word的COM对象提取PDF文本(无需额外安装,Office 2013+默认支持PDF读取),结合哈希集合快速匹配ID,适合处理大量文件:

# 定义路径
$pdfFolder = "C:\temp\pdf"
$idFile = "C:\temp\pdf\ids.txt"
$targetFolder = "C:\temp\pdf\matched_files"

# 创建目标文件夹(如果不存在)
if (-not (Test-Path $targetFolder)) {
    New-Item -ItemType Directory -Path $targetFolder | Out-Null
}

# 读取ID列表到哈希集合,提升查找效率
$ids = New-Object System.Collections.Generic.HashSet[string]
Get-Content $idFile | ForEach-Object {
    $cleanId = $_.Trim()
    if (-not [string]::IsNullOrWhiteSpace($cleanId)) {
        $ids.Add($cleanId) | Out-Null
    }
}

# 初始化Word应用(后台运行)
$word = New-Object -ComObject Word.Application
$word.Visible = $false

# 遍历所有PDF文件
Get-ChildItem -Path $pdfFolder -Filter *.pdf -File | ForEach-Object {
    try {
        # 只读打开PDF文件
        $doc = $word.Documents.Open($_.FullName, $false, $true)
        $pdfText = $doc.Content.Text
        $doc.Close($false) # 不保存直接关闭

        # 检查是否匹配任意ID
        foreach ($id in $ids) {
            if ($pdfText -match [regex]::Escape($id)) {
                Copy-Item -Path $_.FullName -Destination $targetFolder
                Write-Host "Copied: $($_.Name)"
                break # 找到匹配后停止当前文件的ID检查
            }
        }
    }
    catch {
        Write-Warning "处理失败 $($_.Name): $_"
    }
}

# 清理Word对象
$word.Quit()
[System.Runtime.Interopservices.Marshal]::ReleaseComObject($word) | Out-Null

VBScript 解决方案

如果更习惯VBS,以下脚本实现相同功能:

' 定义路径
pdfFolder = "C:\temp\pdf"
idFile = "C:\temp\pdf\ids.txt"
targetFolder = "C:\temp\pdf\matched_files"

Set fso = CreateObject("Scripting.FileSystemObject")

' 创建目标文件夹
If Not fso.FolderExists(targetFolder) Then
    fso.CreateFolder(targetFolder)
End If

' 读取ID到字典
Set ids = CreateObject("Scripting.Dictionary")
Set idStream = fso.OpenTextFile(idFile, 1)
Do Until idStream.AtEndOfStream
    cleanId = Trim(idStream.ReadLine())
    If cleanId <> "" And Not ids.Exists(cleanId) Then
        ids.Add cleanId, True
    End If
Loop
idStream.Close

' 初始化Word后台进程
Set word = CreateObject("Word.Application")
word.Visible = False

' 遍历PDF文件
Set folder = fso.GetFolder(pdfFolder)
For Each file In folder.Files
    If LCase(fso.GetExtensionName(file.Name)) = "pdf" Then
        On Error Resume Next
        Set doc = word.Documents.Open(file.Path, False, True)
        If Err.Number = 0 Then
            pdfText = doc.Content.Text
            doc.Close(False)
            ' 匹配ID
            For Each id In ids.Keys
                If InStr(pdfText, id) > 0 Then
                    fso.CopyFile file.Path, targetFolder & "\", True
                    WScript.Echo "Copied: " & file.Name
                    Exit For
                End If
            Next
        Else
            WScript.Echo "处理失败 " & file.Name & ": " & Err.Description
            Err.Clear
        End If
        On Error GoTo 0
    End If
Next

' 清理资源
word.Quit()
Set word = Nothing
Set fso = Nothing

关键注意事项
  • Word依赖:两个方案都需要安装Microsoft Office(2013及以上版本,支持PDF读取),如果没有Office,可以使用独立的pdftotext.exe(无需安装,直接拷贝到C:\temp\pdf文件夹),替换代码中提取文本的部分:
    • PowerShell替换:$pdfText = & "$pdfFolder\pdftotext.exe" -raw $_.FullName -
    • VBS替换:Set shell = CreateObject("WScript.Shell"): pdfText = shell.Exec("""" & pdfFolder & "\pdftotext.exe"" -raw """ & file.Path & """ -").StdOut.ReadAll()
  • 性能优化:使用哈希集合/字典存储ID,避免每次遍历3000个ID的低效操作;处理20000个PDF耗时较长,建议在空闲时段运行。
  • 异常处理:脚本包含基础错误捕获,可跳过损坏的PDF文件。

内容的提问来源于stack exchange,提问作者Veebster

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 10:42:43