You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java如何在不加载完整文件到内存时拆分合并二进制大文件?

Java二进制文件分块拆分与合并实现

针对大二进制文件(如压缩包)的分块拆分与合并,核心是逐块读写,避免一次性加载整个文件到内存,以下是具体实现:

一、文件拆分方法 List<File> fileSplit(File fullFile, int maxChunkSize)

实现思路

通过缓冲输入流逐块读取原文件内容,每读取maxChunkSize字节就写入一个独立的分块文件,直到原文件读取完毕。分块文件采用原文件名+序号的命名规则(如FullFile.part0),确保后续合并时的顺序正确性。

代码实现

import java.io.*;
import java.util.ArrayList;
import java.util.List;

public class FileSplitterJoiner {
    public List<File> fileSplit(File fullFile, int maxChunkSize) throws IOException {
        List<File> chunkFiles = new ArrayList<>();
        if (!fullFile.exists() || !fullFile.isFile()) {
            throw new IllegalArgumentException("原文件不存在或不是有效文件");
        }
        if (maxChunkSize <= 0) {
            throw new IllegalArgumentException("分块大小必须大于0");
        }

        try (BufferedInputStream bis = new BufferedInputStream(new FileInputStream(fullFile))) {
            byte[] buffer = new byte[maxChunkSize];
            int bytesRead;
            int chunkIndex = 0;

            while ((bytesRead = bis.read(buffer)) != -1) {
                File chunkFile = new File(fullFile.getParent(), fullFile.getName() + ".part" + chunkIndex);
                try (BufferedOutputStream bos = new BufferedOutputStream(new FileOutputStream(chunkFile))) {
                    bos.write(buffer, 0, bytesRead);
                }
                chunkFiles.add(chunkFile);
                chunkIndex++;
            }
        }
        return chunkFiles;
    }
}

二、文件合并方法 File fileJoin(List<File> splitFiles)

实现思路

按传入的分块文件顺序,逐个通过缓冲输入流读取内容并写入目标合并文件。必须保证分块文件的顺序与拆分时一致,否则会导致二进制文件损坏。

代码实现

public File fileJoin(List<File> splitFiles) throws IOException {
    if (splitFiles == null || splitFiles.isEmpty()) {
        throw new IllegalArgumentException("分块文件列表不能为空");
    }
    // 从第一个分块推导原文件名(去除.part后缀)
    File firstChunk = splitFiles.get(0);
    String originalFileName = firstChunk.getName().replaceAll("\\.part\\d+$", "");
    File mergedFile = new File(firstChunk.getParent(), originalFileName);

    try (BufferedOutputStream bos = new BufferedOutputStream(new FileOutputStream(mergedFile))) {
        byte[] buffer = new byte[8192]; // 8KB缓冲,可按需调整
        int bytesRead;

        for (File chunk : splitFiles) {
            if (!chunk.exists() || !chunk.isFile()) {
                throw new IOException("无效分块文件:" + chunk.getAbsolutePath());
            }
            try (BufferedInputStream bis = new BufferedInputStream(new FileInputStream(chunk))) {
                while ((bytesRead = bis.read(buffer)) != -1) {
                    bos.write(buffer, 0, bytesRead);
                }
            }
        }
    }
    return mergedFile;
}

关键注意事项

  • 缓冲流优化:使用BufferedInputStream和BufferedOutputStream减少磁盘IO次数,大幅提升大文件处理效率。
  • 分块大小选择:建议根据云存储的分块限制(多数服务商支持10MB-5GB分块)设置maxChunkSize,同时缓冲数组不宜过大,避免内存溢出。
  • 顺序保障:拆分时的分块命名需包含有序序号,合并时必须严格按拆分顺序传入文件列表,否则会导致文件损坏。
  • 异常处理:完善的异常捕获能避免文件不存在、权限不足等问题导致的程序崩溃,便于问题排查。

内容的提问来源于stack exchange,提问作者Nicholas DiPiazza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 14:45:34