使用PowerShell读取文件部分流并计算哈希值的实现方法
PowerShell 限定文件读取范围计算哈希高性能实现方案
核心思路:直接调用.NET 原生FileStream控制读取范围,完全避免PowerShell内置cmdlet的额外性能开销,性能和全文件流式哈希方案处于同一水平,仅读取目标范围字节,不会加载全文件。
function Get-FilePartialHash { param ( [Parameter(Mandatory=$true)] [string]$FilePath, [int]$HeadBytes = 1048576, # 默认读取头部1MB,单位字节 [int]$TailBytes = 1048576, # 默认读取尾部1MB,单位字节 [string]$Algorithm = "SHA256" ) # 校验参数 if ($HeadBytes -lt 0 -or $TailBytes -lt 0) { throw "读取字节数不能为负数" } if ($HeadBytes -eq 0 -and $TailBytes -eq 0) { throw "头部和尾部读取字节数不能同时为0" } $hash = [Security.Cryptography.HashAlgorithm]::Create($Algorithm) $buffer = New-Object byte[] 4096 # 4KB读取缓冲区,可根据需求调整大小 $bytesRead = 0 try { $stream = [System.IO.File]::OpenRead($FilePath) $fileLength = $stream.Length # 计算头部哈希 if ($HeadBytes -gt 0) { $remaining = [Math]::Min($HeadBytes, $fileLength) while ($remaining -gt 0) { $count = $stream.Read($buffer, 0, [Math]::Min($buffer.Length, $remaining)) if ($count -eq 0) { break } $hash.TransformBlock($buffer, 0, $count, $buffer, 0) | Out-Null $remaining -= $count } } # 计算尾部哈希 if ($TailBytes -gt 0 -and $TailBytes -lt $fileLength) { $stream.Seek(-$TailBytes, [System.IO.SeekOrigin]::End) | Out-Null $remaining = $TailBytes while ($remaining -gt 0) { $count = $stream.Read($buffer, 0, [Math]::Min($buffer.Length, $remaining)) if ($count -eq 0) { break } $hash.TransformBlock($buffer, 0, $count, $buffer, 0) | Out-Null $remaining -= $count } } # 结束哈希计算 $hash.TransformFinalBlock($buffer, 0, 0) | Out-Null $hexHash = -join ($hash.Hash | ForEach-Object { "{0:x2}" -f $_ }) return $hexHash } finally { if ($stream) { $stream.Close() } if ($hash) { $hash.Dispose() } } }
使用说明
- 仅计算文件前1MB的SHA256:
Get-FilePartialHash -FilePath "C:\测试文件.iso" -TailBytes 0 - 仅计算文件后512KB的MD5:
Get-FilePartialHash -FilePath "C:\测试文件.iso" -HeadBytes 0 -TailBytes 524288 -Algorithm MD5 - 同时计算前2MB+后2MB的SHA1:
Get-FilePartialHash -FilePath "C:\测试文件.iso" -HeadBytes 2MB -TailBytes 2MB -Algorithm SHA1
性能说明
原方案2性能差的原因
Get-Content即使限定了读取字节数,其读取的字节流会经过PowerShell管道处理,存在大量的类型装箱拆箱、对象转换开销,导致性能断崖式下跌,直接调用.NET原生IO类可以完全规避该问题。
该方案直接走.NET原生IO操作,无额外转换开销,3GB文件读取前后各1MB的总耗时在100毫秒以内,性能仅取决于磁盘的随机读写速度,远优于Get-Content方案。
内容的提问来源于stack exchange,提问作者Jamesmontalvo3
相关产品推荐
相关产品推荐

