You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Files.copy处理大文件触发OutOfMemoryError的原因及替代方案

解决大文件GZIP压缩时的Java堆内存溢出问题

你的代码核心问题是将整个压缩后的文件内容全部加载到内存(ByteArrayOutputStream + toByteArray()),1GB的原始文件压缩后仍可能占用数百MB内存,远超堆内存限制,直接触发OutOfMemoryError。

以下是几种无需使用IOUtils.copyLarge的替代方案:

方案1:直接写入目标文件(推荐,最简单)

放弃返回byte数组,直接将压缩结果写入磁盘文件,全程流式处理,内存仅占用流的缓冲区大小(默认8KB左右):

private void gzipCompressToFile(File target) throws IOException {
    File source = super.getPath();
    try (GZIPOutputStream gos = new GZIPOutputStream(new FileOutputStream(target))) {
        Files.copy(source.toPath(), gos);
    }
}

方案2:使用NIO Channel优化传输

利用Java NIO的FileChannel.transferTo方法,底层可能采用零拷贝技术,效率更高,内存占用同样极低:

private void gzipCompressWithNio(File target) throws IOException {
    File source = super.getPath();
    try (FileChannel sourceChannel = FileChannel.open(source.toPath(), StandardOpenOption.READ);
         GZIPOutputStream gos = new GZIPOutputStream(new FileOutputStream(target));
         WritableByteChannel targetChannel = Channels.newChannel(gos)) {
        sourceChannel.transferTo(0, sourceChannel.size(), targetChannel);
    }
}

方案3:返回流式InputStream(需返回数据时用)

如果必须返回压缩后的数据而非写入文件,可以返回InputStream,通过管道流在单独线程中异步压缩,调用方可以逐步读取数据,避免内存堆积:

private InputStream gzipCompressAsStream() throws IOException {
    PipedInputStream pis = new PipedInputStream();
    PipedOutputStream pos = new PipedOutputStream(pis);
    File source = super.getPath();

    new Thread(() -> {
        try (GZIPOutputStream gos = new GZIPOutputStream(pos);
             FileInputStream fis = new FileInputStream(source)) {
            byte[] buffer = new byte[8192];
            int readBytes;
            while ((readBytes = fis.read(buffer)) != -1) {
                gos.write(buffer, 0, readBytes);
            }
        } catch (IOException e) {
            try {
                pis.close();
            } catch (IOException ex) {
                ex.printStackTrace();
            }
        }
    }).start();

    return pis;
}

核心思路

所有方案的本质都是避免一次性加载全部数据到内存,采用流式处理模式,让数据通过固定大小的缓冲区逐步完成读写,从根源上解决堆内存溢出问题。

内容的提问来源于stack exchange,提问作者user1472672

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 15:51:52