Camel路由中解压Zip后按扩展名排序文件流的方案咨询
解决Camel路由中Zip文件内部条目按扩展名排序的问题
一、按扩展名排序Zip内部文件的实现方案
方案1:自定义ZipSplitter(推荐,适合大文件)
默认ZipSplitter会按Zip包内的原始顺序返回条目,我们可以继承DefaultZipSplitter,重写prepareEntries方法对ZipEntry按扩展名排序,这样拆分出的条目就是指定顺序,同时保持流式处理的内存友好性,完全适配2GB级别的大Zip文件:
public class SortedZipSplitter extends DefaultZipSplitter { @Override protected List<ZipEntry> prepareEntries(ZipFile zipFile) throws IOException { List<ZipEntry> entries = super.prepareEntries(zipFile); // 自定义排序逻辑:优先处理txt文件,再处理jpg entries.sort((entry1, entry2) -> { String name1 = entry1.getName(); String name2 = entry2.getName(); boolean isTxt1 = name1.toLowerCase().endsWith(".txt"); boolean isTxt2 = name2.toLowerCase().endsWith(".txt"); if (isTxt1 && !isTxt2) return -1; // txt排在前面 if (!isTxt1 && isTxt2) return 1; // txt排在前面 return 0; // 同类型保持原顺序 }); return entries; } }
修改路由代码,使用这个自定义Splitter并保留streaming()优化内存:
fromF("file:zipdir?recursive=false&noop=true&delay=1000&shuffle=true&sortBy=file:modified&include=.*\\zip&moveFailed=errorDir") .split(new SortedZipSplitter()) .streaming() .to("file:destinationdir") .end() .to("file:archiveDir");
这个方案仅对Zip的元数据(条目信息)排序,不加载文件内容到内存,性能和内存占用都很友好,完全适配大体积Zip文件。
方案2:聚合后排序(适合已知条目数量的场景)
如果不想自定义Splitter,可以先把Zip拆分出的所有条目聚合到列表中,排序后再逐个处理。该方式会暂存所有条目对应的Exchange到内存,适合Zip内文件数量少的场景(比如你这里固定两个文件):
fromF("file:zipdir?recursive=false&noop=true&delay=1000&shuffle=true&sortBy=file:modified&include=.*\\zip&moveFailed=errorDir") .split(new ZipSplitter()) // 聚合所有拆分出的文件条目 .aggregate(new ArrayListAggregationStrategy()) .completionSize(2) // 若文件数量不固定,替换为completionFromBatchConsumer() .process(exchange -> { List<Exchange> fileExchanges = exchange.getIn().getBody(List.class); // 按扩展名排序 fileExchanges.sort((e1, e2) -> { String fileName1 = e1.getIn().getHeader(Exchange.FILE_NAME, String.class); String fileName2 = e2.getIn().getHeader(Exchange.FILE_NAME, String.class); boolean isTxt1 = fileName1.toLowerCase().endsWith(".txt"); boolean isTxt2 = fileName2.toLowerCase().endsWith(".txt"); if (isTxt1 && !isTxt2) return -1; if (!isTxt1 && isTxt2) return 1; return 0; }); }) // 重新拆分排序后的列表,逐个输出到目标目录 .split(body()) .to("file:destinationdir") .end() .end() .to("file:archiveDir");
二、关于2GB大Zip文件的处理建议
Camel的ZipSplitter默认支持流式处理(开启streaming()),它基于Java的ZipInputStream实现,不会将整个Zip包加载到内存,只会逐条读取并处理单个ZipEntry,因此完全可以应对2GB级别的大文件,无需担心内存溢出问题。
如果考虑用unzip命令结合FileUtil手动处理,需要注意以下问题:
- 跨平台兼容性:
unzip是Linux/Unix下的命令,Windows系统需要额外安装工具或使用PowerShell替代命令; - 异常处理复杂度:需要手动捕获命令执行失败、超时等情况,处理逻辑比Camel组件繁琐;
- 稳定性:Camel组件已封装成熟的文件处理逻辑,比手动调用外部命令更可靠。
综上,推荐使用自定义SortedZipSplitter的方案,既满足排序需求,又能高效处理大体积Zip文件。
内容的提问来源于stack exchange,提问作者Julia
相关产品推荐
相关产品推荐

