You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java读取权重文件方法执行后内存占用过高未释放问题排查

问题背景

开发神经网络程序时,权重数据存储在长txt文件中,使用|作为数字分隔符解析数值存入数组返回。
调用解析方法时程序内存占用升至1500MB,直到程序终止才释放;不调用该方法时内存仅占700MB。尝试关闭相关资源、排查未回收线程后问题仍未解决,待排查代码如下:

public static int[] findDoub(String fileName) {
    ArrayList<Integer> innt = new ArrayList<Integer>();
    try {
        File file = new File(fileName);
        FileInputStream strem = new FileInputStream(file);
        BufferedInputStream buff = new BufferedInputStream(strem);
        
        int index = buff.read();
        StringBuilder doubus = new StringBuilder("");
        for (int i = 0; index != -1; i++) {
            char a = (char) index;
            if (i > 0) {
                if (a != '|') {
                    doubus.append(a);
                } else {
                    innt.add(Integer.valueOf(doubus.toString()));
                    doubus = new StringBuilder("");
                }
            }
            index = buff.read();
        }
        buff.close();
        strem.close();
        buff = null;
        strem = null;
        
    } catch (IOException e) {
        e.printStackTrace();
    }
    Runtime.getRuntime().gc();
    int[] innnt = new int[innt.size()];
    for (int i = 0; i < innnt.length; i++) {
        innnt[i] = innt.get(i);
    }
    return innnt;
}
内存异常占用根因
  • 核心问题是重复存储带来的内存翻倍+包装类内存开销:代码先用ArrayList<Integer>存储解析结果,每个Integer是引用类型,单对象内存开销是基本类型int的4倍(Integer占16字节,int仅占4字节);后续又额外创建等长的int[]数组做值拷贝,此时ArrayList<Integer>和int[]同时在内存中存在,直接拉高内存峰值。
  • 手动GC调用完全无效:Runtime.getRuntime().gc()仅向JVM发送GC建议,不保证立刻触发回收;且执行该行代码时innt列表还在作用域内属于强引用,就算触发GC也无法回收该对象。
  • 流资源关闭逻辑有缺陷:如果读文件过程中抛出异常,buff.close()和strem.close()不会执行,会导致文件句柄泄漏,但该问题不会带来百MB级别的内存增长。
  • 临时对象过多拉高内存峰值:每次遇到分隔符就新建StringBuilder对象,解析过程中生成大量临时String实例,会进一步抬高内存占用。
  • JVM内存归还机制影响观测结果:JVM回收堆内存后不会立刻把内存归还给操作系统,会预留内存供后续对象分配使用,因此会观测到进程内存居高不下,直到程序退出才释放,这属于正常运行机制。
  • 额外逻辑bug:循环中用i>0跳过了第一个读取的字符,会导致文件开头的第一个数值解析错误。
优化后代码
public static int[] findDoub(String fileName) {
    // 第一次遍历统计数值总个数,提前确定数组长度,避免动态扩容开销
    int totalNum = 0;
    try (BufferedInputStream countStream = new BufferedInputStream(new FileInputStream(fileName))) {
        int c;
        while ((c = countStream.read()) != -1) {
            if (c == '|') totalNum++;
        }
    } catch (IOException e) {
        e.printStackTrace();
        return new int[0];
    }

    int[] result = new int[totalNum];
    int arrIdx = 0;
    // try-with-resources 自动关闭流,无需手动释放
    try (BufferedInputStream buff = new BufferedInputStream(new FileInputStream(fileName))) {
        int index = buff.read();
        StringBuilder numBuilder = new StringBuilder();
        while (index != -1) {
            char a = (char) index;
            if (a != '|') {
                numBuilder.append(a);
            } else {
                // 直接解析为int存入基本类型数组,不生成Integer包装对象
                result[arrIdx++] = Integer.parseInt(numBuilder.toString());
                // 重置StringBuilder,避免重复创建对象
                numBuilder.setLength(0);
            }
            index = buff.read();
        }
    } catch (IOException e) {
        e.printStackTrace();
    }
    return result;
}

优化点说明

  • 移除中间ArrayList<Integer>存储层,直接创建基本类型int[]存储结果,既避免了包装类的额外内存开销,也不需要二次数组拷贝,内存占用可降至原实现的1/4左右。
  • 用StringBuilder.setLength(0)重置缓存,避免每次遇到分隔符都新建StringBuilder对象,减少临时对象生成,降低内存峰值。
  • 用try-with-resources语法自动管理流资源,异常场景下也能正常关闭流,杜绝句柄泄漏。
  • 删除无意义的手动GC调用,JVM会根据运行状态自动触发合适的垃圾回收。
  • 移除多余的循环计数器,修复首字符被跳过的解析bug。

内容的提问来源于stack exchange,提问作者Juan_Puthr

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.28 15:45:49