大文件字节序列处理与拆分的VB分块读取方案问询
数字考古大文件处理:高效分块读取方案需求与实现
我从事数字考古工作,从磁带中恢复了数GB级的大文件。由于所用归档软件老旧且可能为专有软件,我希望编写自定义工具实现两项操作:
- 查找并移除多种备份软件标记类字节序列(含头部、固定长度校验和等内容)
- 通过特定字节标记拆分文件(可识别头部但无法解码,在头部间拆分,拆分后文件从KB到数百MB不等)
此前我使用Microsoft Visual Basic的Read All Bytes方法实现了功能,但该方法存在2GB文件大小限制,请问是否有类似旧系统FileSystem.Seek Method的高效分块读取方法?
临时实现代码
Private Structure Radom Dim startas As List(Of Integer) Dim stopas As List(Of Integer) End Structure Private Function FindIt(ByRef bytes As Byte(), ByRef search As Byte(), limitas As UInt32) As Radom Dim index As Int32 Dim i As UInt32 = 0 Dim l As UInt32 = 0 Dim z As Radom z.startas = New List(Of Integer) z.stopas = New List(Of Integer) While index < limitas - 5 And index >= 0 index = Array.IndexOf(bytes, search(0), index) If index < 0 Then Exit While End If If Confirmdit(bytes, search, index) Then index += search.Length If (i > 0) Then z.startas.Add(l) z.stopas.Add(index - search.Length) End If l = index - search.Length i += 1 End If index += 1 End While If l > 0 Then z.startas.Add(l) z.stopas.Add(limitas - 1) End If Return z End Function Private Function Confirmdit(ByRef buferis As Byte(), ByRef ieskom As Byte(), start As UInt32) As Boolean Dim b As Boolean = True For i = 0 To ieskom.Length - 1 If start + i > buferis.Length - 1 Then b = False Exit For End If If buferis(start + i) <> ieskom(i) Then b = False Exit For End If Next Confirmdit = b End Function
改进后的分块读取实现(已修复问题)
最初实现存在匹配数量偏差问题:在1.7GB的文件中,该方法找到26315处字节序列,而使用整缓冲区读取和HexEdit软件仅找到26156处。现已修复。
Private Function IndexFile(filename As String) As Radom Dim z As Radom Dim k As Radom Dim tmp_c As Integer = 0 Const C_ROLLBACK As UInteger = 200 'must be bigger than searchable bytes z.startas = New List(Of Integer) z.stopas = New List(Of Integer) Dim buferis(&H1000000) As Byte Dim ReadBytes As Int32 Dim currentpos As Int32 = 0 Dim i As Integer Using fsSrc As New FileStream(filename, FileMode.Open, FileAccess.Read, FileShare.Read, &H1000000) While currentpos < fsSrc.Length ReadBytes = fsSrc.Read(buferis, 0, &H1000000) k = FindIt(buferis, magic0, ReadBytes) ' (source buffer, bytes to find, search lenght) RichTextBox1.AppendText(k.startas.Count.ToString + vbCrLf) tmp_c += k.startas.Count For i = 0 To k.startas.Count - 1 If z.startas.BinarySearch(k.startas(i) + currentpos) < 0 Then z.startas.Add(k.startas(i) + currentpos) z.stopas.Add(k.stopas(i) + currentpos) End If Next currentpos += ReadBytes If currentpos < fsSrc.Length Then currentpos = fsSrc.Seek(-C_ROLLBACK, SeekOrigin.Current) End If End While End Using RichTextBox1.AppendText("VISO:" + tmp_c.ToString + vbCrLf) RichTextBox1.AppendText("NVISO:" + z.startas.Count.ToString + vbCrLf) Return z End Function
注意事项
- 处理更大文件时,需将大量
Int32和UInt32类型改为64位值 Array.IndexOf仅支持有符号Int32,这限制了缓冲区大小最大为2GB- 当前文件索引功能运行正常,且速度相当快
内容的提问来源于stack exchange,提问作者Kauno Medis
相关产品推荐
相关产品推荐

