PowerShell如何准确区分文本与二进制文件并提取混合文件文本?
问题描述
我写了一个PowerShell脚本,用来遍历指定文件夹并搜索文件内文本。现在遇到两个问题:
- 纯文本文件误判:比如
C:\csb.log这类系统日志文件被错误识别为二进制文件 - 混合格式文件处理困难:PDF、EPUB/MOBI这类半文本半二进制的文件,无法有效提取其中的文本内容
我不想依赖硬编码的扩展名列表来判断文件类型(如下代码),希望能通过文件内容而非扩展名快速准确识别可直接搜索的纯文本文件,同时解决混合格式文件的文本提取问题。
$binaryExtensions = @('.exe', '.dll', '.bin', '.iso', '.zip', '.tar', '.rar', '.7z', '.gz', '.pdf', '.epub', '.mobi', '.azw', '.azw2', '.azw3')
以下是我目前实现的Is-BinaryFile函数:
function Is-BinaryFile { param ( [string]$FilePath, [switch]$errors ) # Define the log file path based on the function name $functionName = $MyInvocation.MyCommand.Name $logFile = "$Env:TEMP\$functionName.log" # If $errors is specified without $FilePath, output the log file contents if ($errors -and -not $FilePath) { if (Test-Path $logFile) { Get-Content -Path $logFile } else { Write-Host "Log file not found: $logFile" } return } # If the FilePath is a directory, consider it to be a 'binary file' for this and return true if ((Test-Path $FilePath) -and (Get-Item $FilePath).PSIsContainer) { return $true } # Check for common binary file extensions before reading the file $binaryExtensions = @('.exe', '.dll', '.bin', '.iso', '.zip', '.tar', '.rar', '.7z', '.gz', '.pdf', '.epub', '.mobi', '.azw', '.azw2', '.azw3') $fileExtension = [System.IO.Path]::GetExtension($FilePath).ToLower() if ($binaryExtensions -contains $fileExtension) { return $true } # Retry opening the file if it's locked using FileStream in shared mode $maxRetries = 5 $retryDelay = 2 # in seconds $attempt = 0 $fileStream = $null while ($attempt -lt $maxRetries -and -not $fileStream) { try { # Open file with shared read/write mode $fileStream = [System.IO.FileStream]::new($FilePath, [System.IO.FileMode]::Open, [System.IO.FileAccess]::Read, [System.IO.FileShare]::ReadWrite) } catch { Write-Host "File is locked, retrying in $retryDelay seconds..." Start-Sleep -Seconds $retryDelay $attempt++ } } if (-not $fileStream) { $timestamp = Get-Date -Format "yyyy-MM-dd_HH-mm-ss" $logMessage = "[$timestamp] Failed to open '$FilePath' after $maxRetries attempts" Write-Host $logMessage $logMessage | Out-File -FilePath $logFile -Append return $false } # Rest of your code to check if it's a binary file... $reader = $null try { $reader = New-Object System.IO.BinaryReader($fileStream) # Check for BOM (Byte Order Mark) $bomBuffer = New-Object byte[] 4 $bytesRead = $reader.Read($bomBuffer, 0, $bomBuffer.Length) if ($bytesRead -ge 2 -and ( ($bomBuffer[0] -eq 0xEF -and $bomBuffer[1] -eq 0xBB -and $bomBuffer[2] -eq 0xBF) -or # UTF-8 BOM ($bomBuffer[0] -eq 0xFF -and $bomBuffer[1] -eq 0xFE) -or # UTF-16 LE BOM ($bomBuffer[0] -eq 0xFE -and $bomBuffer[1] -eq 0xFF) # UTF-16 BE BOM )) { return $false # It's a text file } # If no BOM, continue checking for non-printable characters $binaryBytes = 0 $textBytes = 0 $buffer = New-Object byte[] 1024 while (($bytesRead = $reader.Read($buffer, 0, $buffer.Length)) -gt 0) { for ($i = 0; $i -lt $bytesRead; $i++) { if ($buffer[$i] -eq 0) { $binaryBytes++ } elseif ($buffer[$i] -lt 32 -and $buffer[$i] -ne 9 -and $buffer[$i] -ne 10 -and $buffer[$i] -ne 13) { $binaryBytes++ } else { $textBytes++ } } } return $binaryBytes -gt $textBytes } finally { if ($reader) { $reader.Close() } if ($fileStream) { $fileStream.Close() } } }
问题分析
当前Is-BinaryFile函数的核心问题在于:
- 纯文本文件误判:仅通过二进制字节数是否超过文本字节数判断,像日志文件这类可能包含少量控制字符(如ESC、DEL)的纯文本文件,会被误判为二进制
- 混合格式文件处理缺失:PDF/EPUB/MOBI属于结构化容器格式,内部包含压缩或编码的文本,直接读二进制无法提取有效内容,且当前逻辑靠扩展名直接排除,不符合“不依赖扩展名”的需求
解决方案
一、改进纯文本文件检测逻辑
调整判断规则,重点优化以下几点:
- 空字节是二进制文件的强特征,检测到直接判定为二进制
- 允许一定比例的不可打印字符(比如≤5%),避免少量控制字符导致误判
- 保留BOM检测逻辑,有文本BOM直接判定为文本文件
改进后的Is-BinaryFile函数:
function Is-BinaryFile { param ( [string]$FilePath, [switch]$errors ) $functionName = $MyInvocation.MyCommand.Name $logFile = "$Env:TEMP\$functionName.log" if ($errors -and -not $FilePath) { return if (Test-Path $logFile) { Get-Content $logFile } else { Write-Host "Log file not found: $logFile" } } if ((Test-Path $FilePath) -and (Get-Item $FilePath).PSIsContainer) { return $true } # 保留文件锁定重试逻辑 $maxRetries = 5 $retryDelay = 2 $attempt = 0 $fileStream = $null while ($attempt -lt $maxRetries -and -not $fileStream) { try { $fileStream = [System.IO.FileStream]::new($FilePath, [System.IO.FileMode]::Open, [System.IO.FileAccess]::Read, [System.IO.FileShare]::ReadWrite) } catch { Write-Host "File is locked, retrying in $retryDelay seconds..." Start-Sleep -Seconds $retryDelay $attempt++ } } if (-not $fileStream) { $timestamp = Get-Date -Format "yyyy-MM-dd_HH-mm-ss" $logMessage = "[$timestamp] Failed to open '$FilePath' after $maxRetries attempts" Write-Host $logMessage $logMessage | Out-File -FilePath $logFile -Append return $false } $reader = $null try { $reader = New-Object System.IO.BinaryReader($fileStream) $bomBuffer = New-Object byte[] 4 $bytesRead = $reader.Read($bomBuffer, 0, $bomBuffer.Length) # BOM检测:有文本BOM直接判定为文本文件 $hasTextBom = $bytesRead -ge 2 -and ( ($bomBuffer[0] -eq 0xEF -and $bomBuffer[1] -eq 0xBB -and $bomBuffer[2] -eq 0xBF) -or ($bomBuffer[0] -eq 0xFF -and $bomBuffer[1] -eq 0xFE) -or ($bomBuffer[0] -eq 0xFE -and $bomBuffer[1] -eq 0xFF) ) if ($hasTextBom) { return $false } # 重置流位置,重新读取全部内容 $fileStream.Position = 0 $buffer = New-Object byte[] 4096 $totalBytes = 0 $nullBytes = 0 $nonPrintableBytes = 0 while (($bytesRead = $reader.Read($buffer, 0, $buffer.Length)) -gt 0) { $totalBytes += $bytesRead for ($i = 0; $i -lt $bytesRead; $i++) { $byte = $buffer[$i] # 空字节是二进制强特征,直接返回true if ($byte -eq 0) { return $true } # 统计不可打印字符(排除制表符、换行、回车) elseif ($byte -lt 32 -and $byte -notin @(9,10,13)) { $nonPrintableBytes++ } } } # 不可打印字符占比超过5%则判定为二进制 if ($totalBytes -eq 0) { return $false } # 空文件视为文本 $nonPrintableRatio = $nonPrintableBytes / $totalBytes return $nonPrintableRatio -gt 0.05 } finally { $reader?.Close() $fileStream?.Close() } }
二、混合格式文件的文本提取
PDF、EPUB/MOBI这类文件属于结构化容器,需要专用工具解析文本内容,以下是PowerShell集成方案:
1. PDF文本提取:使用Poppler的pdftotext
- 下载Poppler工具包,将
pdftotext.exe路径加入系统环境变量 - PowerShell调用示例:
function Extract-PdfText { param([string]$FilePath) $tempTxt = [System.IO.Path]::GetTempFileName() + ".txt" & pdftotext.exe -nopgbrk "$FilePath" "$tempTxt" if ($LASTEXITCODE -eq 0) { $content = Get-Content $tempTxt -Raw Remove-Item $tempTxt -Force return $content } return $null }
2. EPUB/MOBI文本提取:使用Calibre的ebook-convert
- 下载Calibre,将安装目录下的
ebook-convert.exe路径加入系统环境变量 - PowerShell调用示例:
function Extract-EbookText { param([string]$FilePath) $tempTxt = [System.IO.Path]::GetTempFileName() + ".txt" & ebook-convert.exe "$FilePath" "$tempTxt" --no-default-epub-cover if ($LASTEXITCODE -eq 0) { $content = Get-Content $tempTxt -Raw Remove-Item $tempTxt -Force return $content } return $null }
3. 整合搜索逻辑
遍历文件时,先通过改进后的Is-BinaryFile判断:
- 返回
$false:直接读取文件内容进行搜索 - 返回
$true:通过文件签名识别格式(比如PDF开头为%PDF-,EPUB为ZIP容器),调用对应工具提取文本后再搜索
总结
- 纯文本检测通过空字节强特征+不可打印字符比例阈值实现,避免日志类文件的误判
- 混合格式文件依赖专用工具解析,无需硬编码扩展名,通过文件签名识别格式后调用对应提取逻辑
内容的提问来源于stack exchange,提问作者YorSubs
相关产品推荐
相关产品推荐

