You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PowerShell脚本优化:提取PDF关键词及后续字符并导出CSV

Got it, let's fix up your PowerShell script to meet your two main needs: extracting "Address" plus the next 200 characters, and exporting all results to a CSV instead of just printing to the console. Here's a revised script with explanations of the key changes:


Revised PowerShell Script for Address Extraction & CSV Export
# Load the iTextSharp DLL once (more efficient than loading in every loop)
Add-Type -Path "C:\Users\U6013303\Desktop\Muni Refresh\Archives\itextsharp.dll"

# Define paths as variables for easier maintenance
$pdfFolder = "C:\Users\U6013303\Desktop\Muni Refresh\DOC\old\4295479598"
$outputCsv = "C:\Users\U6013303\Desktop\Muni Refresh\Address_Extraction_Results.csv"

# Initialize an array to store our structured results
$extractedResults = @()

# Get all PDF files in the target directory
$pdflist = Get-ChildItem -Path $pdfFolder -Filter "*.pdf"

foreach ($pdff in $pdflist) {
    Write-Host "Processing file: $($pdff.Name)" -BackgroundColor Black -ForegroundColor Green
    $pdfReader = New-Object iTextSharp.text.pdf.pdfreader -ArgumentList $pdff.FullName

    # Loop through each page in the PDF
    for ($page = 1; $page -le $pdfReader.NumberOfPages; $page++) {
        $extractionStrategy = New-Object 'iTextSharp.text.pdf.parser.SimpleTextExtractionStrategy'
        $pageText = [iTextSharp.text.pdf.parser.PdfTextExtractor]::GetTextFromPage($pdfReader, $page, $extractionStrategy)
        
        # Fix encoding issues to ensure consistent text handling
        $cleanPageText = [system.text.Encoding]::UTF8.GetString([System.Text.ASCIIEncoding]::Convert(
            [system.text.encoding]::default, [system.text.encoding]::UTF8, [system.text.Encoding]::Default.GetBytes($pageText)))

        # Regex to capture "Address" + up to 200 following characters (including newlines)
        $addressMatches = [regex]::Matches($cleanPageText, '(Address[\s\S]{0,200})')
        
        # Add each match to our results with context
        foreach ($match in $addressMatches) {
            $extractedResults += [PSCustomObject]@{
                'File Name' = $pdff.Name
                'Page Number' = $page
                'Extracted Address Content' = $match.Value.Trim()
            }
        }
    }

    $pdfReader.Close()
}

# Export results to CSV (UTF8 encoding ensures special characters display correctly)
$extractedResults | Export-Csv -Path $outputCsv -NoTypeInformation -Encoding UTF8

Write-Host "Done! Results saved to: $outputCsv" -BackgroundColor Black -ForegroundColor Green

Key Changes Explained

  • Moved DLL loading outside the loop: Loading itextsharp.dll once at the start is far more efficient than reloading it for every PDF file.
  • Updated regex pattern: (Address[\s\S]{0,200}) uses [\s\S] to match any character (including newlines, which the default . won't capture), so it grabs the full context even if the text wraps across lines in the PDF. {0,200} specifies we want up to 200 characters after "Address".
  • Structured result objects: Each match is stored as a custom object with metadata (file name, page number, extracted content), so your CSV will have clear, usable columns.
  • Reliable file path handling: Using $pdff.FullName avoids manual path construction errors.
  • CSV export with proper encoding: Export-Csv converts the structured objects to a clean CSV, and -Encoding UTF8 ensures special characters aren't corrupted.

内容的提问来源于stack exchange,提问作者Jungle Boy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 21:18:12