PowerShell脚本优化:提取PDF关键词及后续字符并导出CSV
Got it, let's fix up your PowerShell script to meet your two main needs: extracting "Address" plus the next 200 characters, and exporting all results to a CSV instead of just printing to the console. Here's a revised script with explanations of the key changes:
Revised PowerShell Script for Address Extraction & CSV Export
# Load the iTextSharp DLL once (more efficient than loading in every loop) Add-Type -Path "C:\Users\U6013303\Desktop\Muni Refresh\Archives\itextsharp.dll" # Define paths as variables for easier maintenance $pdfFolder = "C:\Users\U6013303\Desktop\Muni Refresh\DOC\old\4295479598" $outputCsv = "C:\Users\U6013303\Desktop\Muni Refresh\Address_Extraction_Results.csv" # Initialize an array to store our structured results $extractedResults = @() # Get all PDF files in the target directory $pdflist = Get-ChildItem -Path $pdfFolder -Filter "*.pdf" foreach ($pdff in $pdflist) { Write-Host "Processing file: $($pdff.Name)" -BackgroundColor Black -ForegroundColor Green $pdfReader = New-Object iTextSharp.text.pdf.pdfreader -ArgumentList $pdff.FullName # Loop through each page in the PDF for ($page = 1; $page -le $pdfReader.NumberOfPages; $page++) { $extractionStrategy = New-Object 'iTextSharp.text.pdf.parser.SimpleTextExtractionStrategy' $pageText = [iTextSharp.text.pdf.parser.PdfTextExtractor]::GetTextFromPage($pdfReader, $page, $extractionStrategy) # Fix encoding issues to ensure consistent text handling $cleanPageText = [system.text.Encoding]::UTF8.GetString([System.Text.ASCIIEncoding]::Convert( [system.text.encoding]::default, [system.text.encoding]::UTF8, [system.text.Encoding]::Default.GetBytes($pageText))) # Regex to capture "Address" + up to 200 following characters (including newlines) $addressMatches = [regex]::Matches($cleanPageText, '(Address[\s\S]{0,200})') # Add each match to our results with context foreach ($match in $addressMatches) { $extractedResults += [PSCustomObject]@{ 'File Name' = $pdff.Name 'Page Number' = $page 'Extracted Address Content' = $match.Value.Trim() } } } $pdfReader.Close() } # Export results to CSV (UTF8 encoding ensures special characters display correctly) $extractedResults | Export-Csv -Path $outputCsv -NoTypeInformation -Encoding UTF8 Write-Host "Done! Results saved to: $outputCsv" -BackgroundColor Black -ForegroundColor Green
Key Changes Explained
- Moved DLL loading outside the loop: Loading
itextsharp.dllonce at the start is far more efficient than reloading it for every PDF file. - Updated regex pattern:
(Address[\s\S]{0,200})uses[\s\S]to match any character (including newlines, which the default.won't capture), so it grabs the full context even if the text wraps across lines in the PDF.{0,200}specifies we want up to 200 characters after "Address". - Structured result objects: Each match is stored as a custom object with metadata (file name, page number, extracted content), so your CSV will have clear, usable columns.
- Reliable file path handling: Using
$pdff.FullNameavoids manual path construction errors. - CSV export with proper encoding:
Export-Csvconverts the structured objects to a clean CSV, and-Encoding UTF8ensures special characters aren't corrupted.
内容的提问来源于stack exchange,提问作者Jungle Boy
相关产品推荐
相关产品推荐

