使用PowerShell合并.tess文件时文本乱码,求助解决方法
问题原因与解决办法
乱码原因
你的.tess文件是UTF-8编码(无BOM),而Windows PowerShell中的Get-Content默认使用系统ANSI编码读取文件,把UTF-8的希腊字符错误解析成ANSI格式,后续Set-Content又用ANSI编码保存,最终导致乱码。
解决办法
方法1:指定UTF-8编码读写
直接修改原命令,明确指定编码为UTF-8:
Get-Content *.tess -Encoding UTF8 | Set-Content greekcorpus.tess -Encoding UTF8
如果需要生成和原文件格式一致的无BOM UTF-8文件,在PowerShell 5.1及以上版本可以用utf8NoBOM编码:
Get-Content *.tess -Encoding UTF8 | Set-Content greekcorpus.tess -Encoding utf8NoBOM
方法2:用.NET类直接操作文件
绕开PowerShell cmdlet的默认编码限制,用.NET的文件操作类确保UTF-8读写:
$outputPath = "greekcorpus.tess" # 清空已有输出文件 if (Test-Path $outputPath) { Remove-Item $outputPath } # 遍历所有.tess文件并追加内容 Get-ChildItem -Filter *.tess | ForEach-Object { $content = [System.IO.File]::ReadAllText($_.FullName, [System.Text.Encoding]::UTF8) [System.IO.File]::AppendAllText($outputPath, $content, [System.Text.Encoding]::UTF8) }
内容的提问来源于stack exchange,提问作者Gabriel Wallace
相关产品推荐
相关产品推荐

