在Apache Camel中实现大文件流式压缩解压缩(不加载全文件到内存)
处理Apache Camel中大文件的流式压缩/解压缩方案
嘿,我完全懂你遇到的痛点——默认的marshal().gzip()和unmarshall().gzip()确实是把整个文件读进内存操作的,对付几百MB以上的大文件肯定会爆内存。下面给你几个可行的解决方案,都是基于流式处理,不用加载全文件到内存:
1. 使用Camel自带的GzipDataFormat开启流式模式
其实Camel自带的GzipDataFormat早就支持流式处理了,只是默认没开启。你只需要手动配置setStreaming(true),就能实现边读边处理:
压缩示例
GzipDataFormat gzipFormat = new GzipDataFormat(); gzipFormat.setStreaming(true); from("file:/input/raw-files?fileName=large-dataset.txt") .marshal(gzipFormat) .to("file:/output/compressed-files?fileName=large-dataset.txt.gz");
解压缩示例
GzipDataFormat gunzipFormat = new GzipDataFormat(); gunzipFormat.setStreaming(true); from("file:/input/compressed-files?fileName=large-dataset.txt.gz") .unmarshal(gunzipFormat) .to("file:/output/raw-files?fileName=large-dataset.txt");
这个配置会让Camel逐块读取文件内容,完全避免一次性加载大文件到内存,完美解决OOM问题。
2. 自定义Processor结合Java原生流式GZIP类
如果自带的格式满足不了你的特殊需求(比如自定义缓冲大小、额外的文件校验逻辑),可以自己写Camel Processor,用Java原生的GZIPInputStream/GZIPOutputStream实现完全可控的流式处理:
自定义压缩Processor
public class StreamingGzipProcessor implements Processor { private static final int BUFFER_SIZE = 16384; // 16KB缓冲块,可按需调整 @Override public void process(Exchange exchange) throws Exception { File inputFile = exchange.getIn().getBody(File.class); String outputPath = "/output/compressed-files/" + inputFile.getName() + ".gz"; File outputFile = new File(outputPath); try (FileInputStream fis = new FileInputStream(inputFile); GZIPOutputStream gzos = new GZIPOutputStream(new FileOutputStream(outputFile), BUFFER_SIZE)) { byte[] buffer = new byte[BUFFER_SIZE]; int bytesRead; while ((bytesRead = fis.read(buffer)) != -1) { gzos.write(buffer, 0, bytesRead); } } exchange.getOut().setBody(outputFile); } }
路由中调用:
from("file:/input/raw-files?fileName=large-dataset.txt") .process(new StreamingGzipProcessor()) .to("file:/output/compressed-files");
自定义解压缩Processor
类似地,解压缩用GZIPInputStream实现:
public class StreamingGunzipProcessor implements Processor { private static final int BUFFER_SIZE = 16384; @Override public void process(Exchange exchange) throws Exception { File inputFile = exchange.getIn().getBody(File.class); String outputFileName = inputFile.getName().replace(".gz", ""); File outputFile = new File("/output/raw-files/" + outputFileName); try (GZIPInputStream gzis = new GZIPInputStream(new FileInputStream(inputFile), BUFFER_SIZE); FileOutputStream fos = new FileOutputStream(outputFile)) { byte[] buffer = new byte[BUFFER_SIZE]; int bytesRead; while ((bytesRead = gzis.read(buffer)) != -1) { fos.write(buffer, 0, bytesRead); } } exchange.getOut().setBody(outputFile); } }
3. 额外注意事项
- 缓冲块大小:上面用的16KB是兼顾内存占用和IO效率的均衡值,你可以根据服务器内存情况调整(比如32KB或8KB)。
- 消息体适配:如果你的路由中消息体不是
File而是InputStream,直接替换代码中的文件流即可,逻辑完全通用。 - Camel版本:
GzipDataFormat的流式配置是在Camel 2.20版本加入的,如果你用的是更旧的版本,建议升级或者直接用自定义Processor方案。
内容的提问来源于stack exchange,提问作者phoenixSid
相关产品推荐
相关产品推荐

