You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Java中Files.walk遍历目录子目录重复读取问题求助

问题描述

实现仅读取指定文件夹子目录的功能时遇到重复读取问题:需求是遍历目标文件夹,找出所有子目录,若子目录包含txt文件则交由DistributionFolder类处理。当前实现对层级较浅的子目录处理正常,但当子目录下存在更深层级子目录(如示例中的1111目录)时,该深层子目录内的内容会被重复读取。

现有代码

主类中拆分目录的代码

private boolean readingDirectory(File files, WriteDatasetInfo writeDatasetInfo) throws IOException, InterruptedException {

    List<Future> futures = new ArrayList<>();

    // 获取可用处理器数量
    int nrThreads = Runtime.getRuntime().availableProcessors();

    // 创建对应数量线程的线程池
    ExecutorService pool = Executors.newFixedThreadPool(nrThreads);

    // 使用Files.walk遍历目录
    try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)))){
        // 过滤出目录,排除默认目录
        paths.filter(Files ::isDirectory)
                .filter(mainReader::isNotTheDefault)
                // 遍历每个目录,交由DistributionFolder处理
                .forEach(
                        path -> {
                            DistributionFolder distributionFolder = new DistributionFolder(this.directory,this.videoNumber);
                            try {
                                System.out.println(path);
                                this.videoNumber=distributionFolder.readingDirectory(path.toFile(),writeDatasetInfo);
                            } catch (IOException e) {
                                throw new RuntimeException(e);
                            } catch (InterruptedException e) {
                                throw new RuntimeException(e);
                            }
                        });

    }catch (Exception e){
        throw new RuntimeException(e);
    }
    finally {
        pool.shutdown();
        while (!pool.awaitTermination(24L, TimeUnit.HOURS)) {
            System.out.println("Not yet. Still waiting for termination");
        }
        Regression.FullRegression(this.directory);
        return true;
    }
}

private static boolean isNotTheDefault(Path path){
    if(!defaultFile.equals(path.toFile())){
        return true;
    }
    return false;
}

DistributionFolder类代码

protected final String directory;
protected int videoNumber;


public DistributionFolder(String directory,int videoNumber){
    this.directory=directory;
    this.videoNumber=videoNumber;
}


protected int readingDirectory(File files, WriteDatasetInfo writeDatasetInfo) throws IOException, InterruptedException {

    List<Future> futures = new ArrayList<>();

    // 获取可用处理器数量
    int nrThreads = Runtime.getRuntime().availableProcessors();

    // 创建对应数量线程的线程池
    ExecutorService pool = Executors.newFixedThreadPool(nrThreads);

    // 使用Files.walk遍历目录下的文件
    try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)))){
        // 过滤出普通文件中的txt文件
        paths.filter(Files::isRegularFile)
                .filter(DistributionFolder::isTXT)
                // 提交任务到线程池处理
                .forEach(
                        path -> {

                            String[] strPath = String.valueOf(path).split("/");

                            pool.submit(new Reader(String.valueOf(path), videoNumber, this.directory, writeDatasetInfo,strPath[strPath.length-2]));
                            this.videoNumber++;

                        });

    }catch (Exception e){
        throw new RuntimeException(e);
    }
    finally {
        // 等待线程池任务完成
        pool.shutdown();
        while (!pool.awaitTermination(24L, TimeUnit.HOURS)) {
            System.out.println("Not yet. Still waiting for termination");
        }
        String[] splitFile = files.toString().split("/");

        MaintainedForRegression.SendMapForRegression(splitFile[splitFile.length-1],this.directory);
        return this.videoNumber;
    }
}

private static boolean isTXT(Path path) {
    String pathFile = String.valueOf(path);
    String[] slittedpath = pathFile.split("/");
    String[] splittedname = slittedpath[slittedpath.length - 1].split("\\.");
    return splittedname[splittedname.length - 1].equals("txt");
}

问题根源

重复读取的核心原因是遍历逻辑重叠:

  • 主类的readingDirectory方法使用Files.walk会递归遍历目标文件夹下所有层级的子目录,每个子目录都会被传给DistributionFolder处理。
  • 而DistributionFolder的readingDirectory方法同样使用Files.walk,会递归遍历当前传入目录下所有层级的txt文件。
  • 例如存在目录A/B/C,主类会遍历到A/B和A/B/C两个目录:处理A/B时会读取到C下的txt文件,处理A/B/C时又会读取一次该文件,导致重复。

解决建议

根据需求提供两种针对性修改方案:

方案一:主类仅遍历直接子目录,由DistributionFolder处理所有层级

如果需求是处理目标文件夹下所有子目录中的txt文件(每个文件仅处理一次),修改主类的遍历逻辑,只获取目标文件夹的直接子目录,不再递归遍历深层目录:

// 替换原来的Files.walk为Files.list,仅遍历当前目录的直接子项
try (Stream<Path> paths = Files.list(Paths.get(String.valueOf(files)))){
    paths.filter(Files ::isDirectory)
            .filter(mainReader::isNotTheDefault)
            .forEach(/* 原有处理逻辑不变 */);
}

这样主类只会把目标文件夹的一级子目录传给DistributionFolder,由它递归处理所有深层的txt文件,每个文件仅被处理一次。

方案二:DistributionFolder仅处理当前目录的直接txt文件

如果需求是每个子目录仅处理自身直接包含的txt文件(不处理子目录中的文件),修改DistributionFolder的遍历逻辑,限制遍历深度为1:

// 添加深度参数1,仅遍历当前目录的直接子项
try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)), 1)){
    paths.filter(Files::isRegularFile)
            .filter(DistributionFolder::isTXT)
            .forEach(/* 原有处理逻辑不变 */);
}

这样每个子目录仅处理自己直接的txt文件,主类遍历所有层级的子目录时也不会出现重复读取的问题。

额外优化建议

  1. 简化isTXT方法,不用拆分路径:
private static boolean isTXT(Path path) {
    return "txt".equals(Files.getFileExtension(path));
}
  1. 线程池可复用,避免每次调用方法都新建,减少资源浪费。

内容的提问来源于stack exchange,提问作者ilie alexandru

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 00:00:13