You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Camel路由中解压Zip后按扩展名排序文件流的方案咨询

解决Camel路由中Zip文件内部条目按扩展名排序的问题

一、按扩展名排序Zip内部文件的实现方案

方案1:自定义ZipSplitter(推荐,适合大文件)

默认ZipSplitter会按Zip包内的原始顺序返回条目,我们可以继承DefaultZipSplitter,重写prepareEntries方法对ZipEntry按扩展名排序,这样拆分出的条目就是指定顺序,同时保持流式处理的内存友好性,完全适配2GB级别的大Zip文件:

public class SortedZipSplitter extends DefaultZipSplitter {
    @Override
    protected List<ZipEntry> prepareEntries(ZipFile zipFile) throws IOException {
        List<ZipEntry> entries = super.prepareEntries(zipFile);
        // 自定义排序逻辑:优先处理txt文件,再处理jpg
        entries.sort((entry1, entry2) -> {
            String name1 = entry1.getName();
            String name2 = entry2.getName();
            
            boolean isTxt1 = name1.toLowerCase().endsWith(".txt");
            boolean isTxt2 = name2.toLowerCase().endsWith(".txt");
            
            if (isTxt1 && !isTxt2) return -1; // txt排在前面
            if (!isTxt1 && isTxt2) return 1;  // txt排在前面
            return 0; // 同类型保持原顺序
        });
        return entries;
    }
}

修改路由代码,使用这个自定义Splitter并保留streaming()优化内存:

fromF("file:zipdir?recursive=false&noop=true&delay=1000&shuffle=true&sortBy=file:modified&include=.*\\zip&moveFailed=errorDir")
    .split(new SortedZipSplitter())
    .streaming()
    .to("file:destinationdir")
    .end()
.to("file:archiveDir");

这个方案仅对Zip的元数据(条目信息)排序,不加载文件内容到内存,性能和内存占用都很友好,完全适配大体积Zip文件。

方案2:聚合后排序(适合已知条目数量的场景)

如果不想自定义Splitter,可以先把Zip拆分出的所有条目聚合到列表中,排序后再逐个处理。该方式会暂存所有条目对应的Exchange到内存,适合Zip内文件数量少的场景(比如你这里固定两个文件):

fromF("file:zipdir?recursive=false&noop=true&delay=1000&shuffle=true&sortBy=file:modified&include=.*\\zip&moveFailed=errorDir")
    .split(new ZipSplitter())
    // 聚合所有拆分出的文件条目
    .aggregate(new ArrayListAggregationStrategy())
        .completionSize(2) // 若文件数量不固定,替换为completionFromBatchConsumer()
        .process(exchange -> {
            List<Exchange> fileExchanges = exchange.getIn().getBody(List.class);
            // 按扩展名排序
            fileExchanges.sort((e1, e2) -> {
                String fileName1 = e1.getIn().getHeader(Exchange.FILE_NAME, String.class);
                String fileName2 = e2.getIn().getHeader(Exchange.FILE_NAME, String.class);
                
                boolean isTxt1 = fileName1.toLowerCase().endsWith(".txt");
                boolean isTxt2 = fileName2.toLowerCase().endsWith(".txt");
                
                if (isTxt1 && !isTxt2) return -1;
                if (!isTxt1 && isTxt2) return 1;
                return 0;
            });
        })
    // 重新拆分排序后的列表,逐个输出到目标目录
    .split(body())
    .to("file:destinationdir")
    .end()
    .end()
.to("file:archiveDir");

二、关于2GB大Zip文件的处理建议

Camel的ZipSplitter默认支持流式处理(开启streaming()),它基于Java的ZipInputStream实现,不会将整个Zip包加载到内存,只会逐条读取并处理单个ZipEntry,因此完全可以应对2GB级别的大文件,无需担心内存溢出问题。

如果考虑用unzip命令结合FileUtil手动处理,需要注意以下问题:

  • 跨平台兼容性:unzip是Linux/Unix下的命令,Windows系统需要额外安装工具或使用PowerShell替代命令;
  • 异常处理复杂度:需要手动捕获命令执行失败、超时等情况,处理逻辑比Camel组件繁琐;
  • 稳定性:Camel组件已封装成熟的文件处理逻辑,比手动调用外部命令更可靠。

综上,推荐使用自定义SortedZipSplitter的方案,既满足排序需求,又能高效处理大体积Zip文件。

内容的提问来源于stack exchange,提问作者Julia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 03:21:14