PowerShell中StreamWriter编码问题与大文件替换优化技术问询
Background
You're tackling a find-replace task for arbitrary files (up to 500MB) with diverse encodings, and the standard Get-Content/Add-Content approach is too memory-heavy and slow. You've already successfully detected the source file's encoding using:
$reader = [IO.StreamReader]::New($sourceFile, $true) $reader.Peek() > $null $encoding = $reader.CurrentEncoding
This gives you a type name like System.Text.UTF8Encoding, and you've hit three key questions along with some valuable testing progress. Let's walk through each one clearly:
Question 1: Converting Encoding Type Names to Usable Instances
You asked if there’s a built-in way to turn those System.Text.*Encoding type strings into working instances, or if you need to build a custom hash table.
First, your initial test with New-Object $reader.CurrentEncoding $False failed for most encodings because:
- The
$Falseparameter was being misread as a constructor argument, not the-Strictflag for COM objects (and onlyUTF8Encodingworks in that COM context).
As you discovered, the correct approach is using the -TypeName parameter with New-Object—this works for all standard text encodings because it properly resolves the type and uses its default constructor (which matches the settings of the encoding detected by StreamReader):
# Works reliably for all standard .NET text encodings $encoding = New-Object -TypeName $reader.CurrentEncoding
No custom hash table is needed here; .NET can resolve the type name directly when you use -TypeName. For encodings that require constructor arguments (like specifying BOM usage), you can pass those after the type name, but since you’re mirroring the source file’s encoding, the default constructor is safe.
Question 2: StreamWriter Not Respecting Specified Encoding
You noticed that when trying Encoding.UTF16, the output file still ended up as UTF8. This almost certainly happened because you weren’t correctly passing the encoding instance to the StreamWriter constructor initially.
Unlike StreamReader, which needs a read operation (like Peek()) to detect encoding, StreamWriter uses the encoding you pass to its constructor immediately. If you omit the encoding parameter, it defaults to UTF-8 without a BOM—that’s why your test output stayed UTF8 even when you intended UTF16.
Your updated constructor call is spot-on:
$writer = [IO.StreamWriter]::New($destinationFile, $true, (New-Object $reader.CurrentEncoding))
The second $true parameter keeps the stream open (handy for appending or continuous writes), and the third parameter ensures the writer uses the exact encoding detected from the source file.
Question 3: Encoding-Specific File Endings
You mentioned some encodings require specific endings (like writing "rn" for UTF8 to match the original file size). Here’s what you need to know to handle this correctly:
Line Endings: Line ending styles are mostly independent of encoding, but they affect file size:
- Windows-style: Carriage Return + Line Feed (
\r\nor "rn" in PowerShell) - Unix-style: Line Feed only (
\nor "`n") - Old Mac-style: Carriage Return only (
\r)
Instead of hardcoding endings, preserve the exact line endings from the source file to avoid size mismatches.
- Windows-style: Carriage Return + Line Feed (
Byte Order Marks (BOMs): Some encodings (like UTF-8 with BOM, UTF-16, UTF-32) use a BOM at the file’s start. When you create a
StreamWriterwith the detected encoding, it will automatically write the correct BOM if the encoding uses one—you can check the BOM bytes with$encoding.GetPreamble()if needed.Size Mismatch Fixes: If your output file size doesn’t match the original, it’s usually due to:
- Incorrect line ending conversion (adding/removing bytes per line)
- Missing or extra BOM bytes
- Accidental character encoding conversion (e.g., a character taking more bytes in the target encoding)
For detailed encoding behavior, refer to the .NET System.Text.Encoding class docs—they outline BOM usage, default line endings, and other encoding-specific nuances you can leverage in your script.
Final Working Approach Recap
Your current implementation is efficient and reliable for large files:
- Detect the source file encoding with
StreamReader(usingPeek()to trigger detection) - Create a matching encoding instance with
New-Object -TypeName $reader.CurrentEncoding - Use
StreamWriterwith the detected encoding to write incrementally (avoiding full file memory loads)
内容的提问来源于stack exchange,提问作者Gordon

