You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何避免Collection数组对象OOM?限制Apache Commons IO listFiles返回集合大小

解决FileUtils.listFiles百万文件OOM问题:限制返回结果数量

Apache Commons IO的FileUtils.listFiles会一次性把所有匹配文件加载到内存集合中,面对数百万级别的文件必然触发OOM。要实现类似SQL LIMIT的分批获取效果,必须换用迭代式遍历方案,而非一次性收集全部结果。

方案1:用FileUtils.iterateFiles手动控制结果数量

iterateFiles返回Iterator<File>,不会提前加载所有文件到内存,你可以手动控制只取前N个文件:

// 初始化NFS目录和正则过滤器
File nfsDir = new File("/your/nfs/target-dir");
IOFileFilter regexFilter = new RegexFileFilter("^your-regex-pattern$");

// 迭代器模式遍历文件(第三个参数为null则不递归子目录,需递归传FileFilterUtils.directoryFileFilter())
Iterator<File> fileIterator = FileUtils.iterateFiles(nfsDir, regexFilter, null);

final int BATCH_SIZE = 1000; // 每次取1000个文件
List<File> batchFiles = new ArrayList<>(BATCH_SIZE);
int currentCount = 0;

// 只取前BATCH_SIZE个文件
while (fileIterator.hasNext() && currentCount < BATCH_SIZE) {
    batchFiles.add(fileIterator.next());
    currentCount++;
}

// 处理这批文件:比如删除指定日期之前的文件
for (File file : batchFiles) {
    if (isFileOlderThanTargetDate(file, targetDate)) {
        file.delete();
    }
}

方案2:自定义带计数限制的过滤器

如果你想让过滤器自动终止匹配,可以写一个自定义IOFileFilter,达到指定数量后不再匹配文件:

public class LimitedCountFileFilter implements IOFileFilter {
    private final IOFileFilter delegateFilter;
    private final int maxCount;
    private int matchedCount = 0;

    public LimitedCountFileFilter(IOFileFilter delegateFilter, int maxCount) {
        this.delegateFilter = delegateFilter;
        this.maxCount = maxCount;
    }

    @Override
    public boolean accept(File file) {
        // 达到数量上限直接返回false
        if (matchedCount >= maxCount) {
            return false;
        }
        // 符合原始过滤条件则计数+1并返回true
        if (delegateFilter.accept(file)) {
            matchedCount++;
            return true;
        }
        return false;
    }

    @Override
    public boolean accept(File dir, String name) {
        // 目录过滤正常处理,保证子目录能被遍历
        return delegateFilter.accept(dir, name);
    }
}

使用示例:

IOFileFilter limitedFilter = new LimitedCountFileFilter(regexFilter, 1000);
Iterator<File> fileIterator = FileUtils.iterateFiles(nfsDir, limitedFilter, FileUtils.directoryFileFilter());

List<File> batchFiles = new ArrayList<>();
fileIterator.forEachRemaining(batchFiles::add); // 最多获取1000个文件

关键说明

  • 每次处理完一批并删除旧文件后,下一次调用iterateFiles会重新遍历NFS目录(此时文件已减少,结果集自然缩小),完全适配你循环调用的需求。
  • 迭代式遍历的内存占用极低,仅保留当前批次的文件对象,彻底避免OOM问题。

内容的提问来源于stack exchange,提问作者Ashish Ranjan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 06:56:54