You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Apache Camel中实现大文件流式压缩解压缩(不加载全文件到内存)

处理Apache Camel中大文件的流式压缩/解压缩方案

嘿,我完全懂你遇到的痛点——默认的marshal().gzip()和unmarshall().gzip()确实是把整个文件读进内存操作的,对付几百MB以上的大文件肯定会爆内存。下面给你几个可行的解决方案,都是基于流式处理,不用加载全文件到内存:

1. 使用Camel自带的GzipDataFormat开启流式模式

其实Camel自带的GzipDataFormat早就支持流式处理了,只是默认没开启。你只需要手动配置setStreaming(true),就能实现边读边处理:

压缩示例

GzipDataFormat gzipFormat = new GzipDataFormat();
gzipFormat.setStreaming(true);

from("file:/input/raw-files?fileName=large-dataset.txt")
    .marshal(gzipFormat)
    .to("file:/output/compressed-files?fileName=large-dataset.txt.gz");

解压缩示例

GzipDataFormat gunzipFormat = new GzipDataFormat();
gunzipFormat.setStreaming(true);

from("file:/input/compressed-files?fileName=large-dataset.txt.gz")
    .unmarshal(gunzipFormat)
    .to("file:/output/raw-files?fileName=large-dataset.txt");

这个配置会让Camel逐块读取文件内容,完全避免一次性加载大文件到内存,完美解决OOM问题。

2. 自定义Processor结合Java原生流式GZIP类

如果自带的格式满足不了你的特殊需求(比如自定义缓冲大小、额外的文件校验逻辑),可以自己写Camel Processor,用Java原生的GZIPInputStream/GZIPOutputStream实现完全可控的流式处理:

自定义压缩Processor

public class StreamingGzipProcessor implements Processor {
    private static final int BUFFER_SIZE = 16384; // 16KB缓冲块,可按需调整

    @Override
    public void process(Exchange exchange) throws Exception {
        File inputFile = exchange.getIn().getBody(File.class);
        String outputPath = "/output/compressed-files/" + inputFile.getName() + ".gz";
        File outputFile = new File(outputPath);

        try (FileInputStream fis = new FileInputStream(inputFile);
             GZIPOutputStream gzos = new GZIPOutputStream(new FileOutputStream(outputFile), BUFFER_SIZE)) {
            byte[] buffer = new byte[BUFFER_SIZE];
            int bytesRead;
            while ((bytesRead = fis.read(buffer)) != -1) {
                gzos.write(buffer, 0, bytesRead);
            }
        }
        exchange.getOut().setBody(outputFile);
    }
}

路由中调用:

from("file:/input/raw-files?fileName=large-dataset.txt")
    .process(new StreamingGzipProcessor())
    .to("file:/output/compressed-files");

自定义解压缩Processor

类似地,解压缩用GZIPInputStream实现:

public class StreamingGunzipProcessor implements Processor {
    private static final int BUFFER_SIZE = 16384;

    @Override
    public void process(Exchange exchange) throws Exception {
        File inputFile = exchange.getIn().getBody(File.class);
        String outputFileName = inputFile.getName().replace(".gz", "");
        File outputFile = new File("/output/raw-files/" + outputFileName);

        try (GZIPInputStream gzis = new GZIPInputStream(new FileInputStream(inputFile), BUFFER_SIZE);
             FileOutputStream fos = new FileOutputStream(outputFile)) {
            byte[] buffer = new byte[BUFFER_SIZE];
            int bytesRead;
            while ((bytesRead = gzis.read(buffer)) != -1) {
                fos.write(buffer, 0, bytesRead);
            }
        }
        exchange.getOut().setBody(outputFile);
    }
}

3. 额外注意事项

  • 缓冲块大小:上面用的16KB是兼顾内存占用和IO效率的均衡值,你可以根据服务器内存情况调整(比如32KB或8KB)。
  • 消息体适配:如果你的路由中消息体不是File而是InputStream,直接替换代码中的文件流即可,逻辑完全通用。
  • Camel版本:GzipDataFormat的流式配置是在Camel 2.20版本加入的,如果你用的是更旧的版本,建议升级或者直接用自定义Processor方案。

内容的提问来源于stack exchange,提问作者phoenixSid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:20:08