You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

读取大压缩文件触发OutOfMemoryException问题求助

解决大压缩文件读取时的OutOfMemoryException问题

从异常堆栈跟踪能明确看到,问题根源是StreamReader.ReadToEnd()把整个10GB+的文件内容一次性加载到内存字符串中,直接导致内存耗尽。你的代码还存在几个逻辑问题,会进一步加重内存负担,以下是具体的修复方案:

核心修复:避免一次性加载整个文件

不要用ReadToEnd()一次性读取全部内容,改成逐行或按固定缓冲区读取,同时逐段匹配正则表达式,每次只在内存中保留一小部分数据。

修复代码逻辑错误

  • 同时使用foreach和while遍历压缩包条目,会导致重复枚举,直接用单个for循环即可。
  • 手动调用zipToOpen.Close()和reader.Close()完全多余,using块会自动释放资源,提前关闭还会导致后续操作出错。
  • 不要反复调用archive.Entries.Count(),把条目总数存成变量,避免重复计算。

优化内存占用

  • 用C# 7.0+支持的ValueTuple替代Tuple,减少内存开销。
  • 如果不需要保留所有报告记录到最后,可以分批生成报告后清理列表,避免内存堆积。

修改后的代码示例

using System.IO;
using System.IO.Compression;
using System.Text.RegularExpressions;
using System.Collections.Generic;

// ... 其他命名空间

using (FileStream zipToOpen = new FileStream(Path.Combine(fileLocation, $"{zipfile}.zip"), FileMode.Open))
using (ZipArchive archive = new ZipArchive(zipToOpen, ZipArchiveMode.Read))
{
    var reportFiles = new List<(string FullName, int TranSetCount, int TranCount)>();
    int totalTypeTrans = 0;
    int totalTypeTranSet = 0;
    int entryCount = archive.Entries.Count;

    for (int index = 0; index < entryCount; index++)
    {
        var entry = archive.Entries[index];
        if (!entry.FullName.StartsWith("asdf"))
            continue;

        backgroundWorker.ReportProgress(index, entryCount);

        int fileTranSet = 0;
        int fileTran = 0;

        // 逐行读取文件,逐行匹配正则,内存仅保留当前行内容
        using (StreamReader reader = new StreamReader(entry.Open()))
        {
            string line;
            while ((line = reader.ReadLine()) != null)
            {
                fileTranSet += Regex.Matches(line, transsetString).Count;
                fileTran += Regex.Matches(line, transstring).Count;
            }
        }

        reportFiles.Add((entry.FullName, fileTranSet, fileTran));
        totalTypeTrans += fileTran;
        totalTypeTranSet += fileTranSet;

        // 可选:如果条目数量极大,分批生成报告并清理列表
        // if (reportFiles.Count > 1000)
        // {
        //     Directory.CreateDirectory(baseReportDirectoryLocation);
        //     createReports(ReportType, reportFiles, totalTypeTranSet, totalTypeTrans);
        //     reportFiles.Clear();
        // }
    }

    // 生成最终报告
    Directory.CreateDirectory(baseReportDirectoryLocation);
    createReports(ReportType, reportFiles, totalTypeTranSet, totalTypeTrans);
}

代码改进说明

  • 用Path.Combine拼接路径,避免手动拼接斜杠的错误。
  • 逐行读取逻辑确保内存占用始终处于极低水平,支持任意大小的文件。
  • 移除了所有不必要的手动资源释放操作,交给using块自动处理。
  • ValueTuple比传统Tuple更轻量,语法也更简洁。

内容的提问来源于stack exchange,提问作者Cleardragonf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:31:00