如何使用PowerShell检测文件文本编码?以C:\Windows下日志文件为例
使用PowerShell检测文件文本编码的方法
首先,PowerShell并没有内置的Get-FileEncoding cmdlet,但可以通过自定义函数或直接调用.NET类两种方式实现编码检测。
一、自定义Get-FileEncoding函数
你可以在PowerShell中定义一个专用函数来检测文件编码,这是最常用的解决方案:
function Get-FileEncoding { param( [Parameter(Mandatory=$True, ValueFromPipelineByPropertyName=$True)] [string]$Path ) $bytes = Get-Content -Path $Path -Encoding Byte -ReadCount 4 -TotalCount 4 if ($bytes.Count -ge 2) { # 检测UTF-16 LE BOM if ($bytes[0] -eq 0xff -and $bytes[1] -eq 0xfe) { return 'Unicode (UTF-16 LE)' } # 检测UTF-16 BE BOM if ($bytes[0] -eq 0xfe -and $bytes[1] -eq 0xff) { return 'Unicode (UTF-16 BE)' } if ($bytes.Count -ge 3) { # 检测UTF-8 BOM if ($bytes[0] -eq 0xef -and $bytes[1] -eq 0xbb -and $bytes[2] -eq 0xbf) { return 'UTF-8 with BOM' } # 检测UTF-7 BOM if ($bytes[0] -eq 0x2b -and $bytes[1] -eq 0x2f -and $bytes[2] -eq 0x76) { return 'UTF-7' } } if ($bytes.Count -ge 4) { # 检测UTF-32 LE BOM if ($bytes[0] -eq 0xff -and $bytes[1] -eq 0xfe -and $bytes[2] -eq 0x00 -and $bytes[3] -eq 0x00) { return 'UTF-32 LE' } # 检测UTF-32 BE BOM if ($bytes[0] -eq 0x00 -and $bytes[1] -eq 0x00 -and $bytes[2] -eq 0xfe -and $bytes[3] -eq 0xff) { return 'UTF-32 BE' } } } # 无BOM的情况,尝试用UTF-8读取判断 try { $null = Get-Content -Path $Path -Encoding UTF8 -ErrorAction Stop return 'UTF-8 (无BOM)' } catch { return '系统默认编码 (ANSI)' } }
定义好函数后,直接结合你已有的文件列表命令使用即可:
Get-ChildItem -Filter '*.log' -Path "C:\Windows" -Recurse -ErrorAction SilentlyContinue | Select-Object FullName, @{Name='Encoding'; Expression={Get-FileEncoding -Path $_.FullName}}
二、直接调用.NET类检测
如果不想自定义函数,也可以直接借助.NET的文件流和编码类实现检测,示例代码如下:
Get-ChildItem -Filter '*.log' -Path "C:\Windows" -Recurse -ErrorAction SilentlyContinue | ForEach-Object { $fs = New-Object System.IO.FileStream($_.FullName, [System.IO.FileMode]::Open, [System.IO.FileAccess]::Read) $encoding = '' try { # 尝试检测UTF-8 BOM $detector = New-Object System.Text.UTF8Encoding($True, $True) $null = $detector.GetChars($fs.ReadByte(), $fs.ReadByte(), $fs.ReadByte(), $fs.ReadByte()) $encoding = 'UTF-8 with BOM' } catch [System.Text.DecoderFallbackException] { $fs.Position = 0 $bomBytes = New-Object byte[] 4 $readCount = $fs.Read($bomBytes, 0, 4) if ($readCount -ge 2) { switch -Wildcard ($bomBytes) { "0xff,0xfe,*" { $encoding = if ($readCount -ge 4 -and $bomBytes[2] -eq 0 -and $bomBytes[3] -eq 0) { 'UTF-32 LE' } else { 'UTF-16 LE' } break } "0xfe,0xff,*" { $encoding = 'UTF-16 BE' break } "0x2b,0x2f,0x76,*" { $encoding = 'UTF-7' break } "0x00,0x00,0xfe,0xff" { $encoding = 'UTF-32 BE' break } default { $encoding = 'UTF-8 (无BOM)或系统默认编码' } } } else { $encoding = 'UTF-8 (无BOM)或系统默认编码' } } finally { $fs.Close() } # 输出结果对象 [PSCustomObject]@{ FullName = $_.FullName Encoding = $encoding } }
注意事项
- 无BOM的编码检测存在局限性,无法完全精准区分无BOM UTF-8和部分系统编码,上述方法通过尝试UTF-8读取做简易判断。
- 两种方法都只读取文件前几个字节,不会加载整个文件,对大文件也有较好的性能表现。
内容的提问来源于stack exchange,提问作者JamesThomasMoon
相关产品推荐
相关产品推荐

