Java中Files.walk遍历目录子目录重复读取问题求助
问题描述
实现仅读取指定文件夹子目录的功能时遇到重复读取问题:需求是遍历目标文件夹,找出所有子目录,若子目录包含txt文件则交由DistributionFolder类处理。当前实现对层级较浅的子目录处理正常,但当子目录下存在更深层级子目录(如示例中的1111目录)时,该深层子目录内的内容会被重复读取。
现有代码
主类中拆分目录的代码
private boolean readingDirectory(File files, WriteDatasetInfo writeDatasetInfo) throws IOException, InterruptedException { List<Future> futures = new ArrayList<>(); // 获取可用处理器数量 int nrThreads = Runtime.getRuntime().availableProcessors(); // 创建对应数量线程的线程池 ExecutorService pool = Executors.newFixedThreadPool(nrThreads); // 使用Files.walk遍历目录 try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)))){ // 过滤出目录,排除默认目录 paths.filter(Files ::isDirectory) .filter(mainReader::isNotTheDefault) // 遍历每个目录,交由DistributionFolder处理 .forEach( path -> { DistributionFolder distributionFolder = new DistributionFolder(this.directory,this.videoNumber); try { System.out.println(path); this.videoNumber=distributionFolder.readingDirectory(path.toFile(),writeDatasetInfo); } catch (IOException e) { throw new RuntimeException(e); } catch (InterruptedException e) { throw new RuntimeException(e); } }); }catch (Exception e){ throw new RuntimeException(e); } finally { pool.shutdown(); while (!pool.awaitTermination(24L, TimeUnit.HOURS)) { System.out.println("Not yet. Still waiting for termination"); } Regression.FullRegression(this.directory); return true; } } private static boolean isNotTheDefault(Path path){ if(!defaultFile.equals(path.toFile())){ return true; } return false; }
DistributionFolder类代码
protected final String directory; protected int videoNumber; public DistributionFolder(String directory,int videoNumber){ this.directory=directory; this.videoNumber=videoNumber; } protected int readingDirectory(File files, WriteDatasetInfo writeDatasetInfo) throws IOException, InterruptedException { List<Future> futures = new ArrayList<>(); // 获取可用处理器数量 int nrThreads = Runtime.getRuntime().availableProcessors(); // 创建对应数量线程的线程池 ExecutorService pool = Executors.newFixedThreadPool(nrThreads); // 使用Files.walk遍历目录下的文件 try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)))){ // 过滤出普通文件中的txt文件 paths.filter(Files::isRegularFile) .filter(DistributionFolder::isTXT) // 提交任务到线程池处理 .forEach( path -> { String[] strPath = String.valueOf(path).split("/"); pool.submit(new Reader(String.valueOf(path), videoNumber, this.directory, writeDatasetInfo,strPath[strPath.length-2])); this.videoNumber++; }); }catch (Exception e){ throw new RuntimeException(e); } finally { // 等待线程池任务完成 pool.shutdown(); while (!pool.awaitTermination(24L, TimeUnit.HOURS)) { System.out.println("Not yet. Still waiting for termination"); } String[] splitFile = files.toString().split("/"); MaintainedForRegression.SendMapForRegression(splitFile[splitFile.length-1],this.directory); return this.videoNumber; } } private static boolean isTXT(Path path) { String pathFile = String.valueOf(path); String[] slittedpath = pathFile.split("/"); String[] splittedname = slittedpath[slittedpath.length - 1].split("\\."); return splittedname[splittedname.length - 1].equals("txt"); }
问题根源
重复读取的核心原因是遍历逻辑重叠:
- 主类的
readingDirectory方法使用Files.walk会递归遍历目标文件夹下所有层级的子目录,每个子目录都会被传给DistributionFolder处理。 - 而
DistributionFolder的readingDirectory方法同样使用Files.walk,会递归遍历当前传入目录下所有层级的txt文件。 - 例如存在目录
A/B/C,主类会遍历到A/B和A/B/C两个目录:处理A/B时会读取到C下的txt文件,处理A/B/C时又会读取一次该文件,导致重复。
解决建议
根据需求提供两种针对性修改方案:
方案一:主类仅遍历直接子目录,由DistributionFolder处理所有层级
如果需求是处理目标文件夹下所有子目录中的txt文件(每个文件仅处理一次),修改主类的遍历逻辑,只获取目标文件夹的直接子目录,不再递归遍历深层目录:
// 替换原来的Files.walk为Files.list,仅遍历当前目录的直接子项 try (Stream<Path> paths = Files.list(Paths.get(String.valueOf(files)))){ paths.filter(Files ::isDirectory) .filter(mainReader::isNotTheDefault) .forEach(/* 原有处理逻辑不变 */); }
这样主类只会把目标文件夹的一级子目录传给DistributionFolder,由它递归处理所有深层的txt文件,每个文件仅被处理一次。
方案二:DistributionFolder仅处理当前目录的直接txt文件
如果需求是每个子目录仅处理自身直接包含的txt文件(不处理子目录中的文件),修改DistributionFolder的遍历逻辑,限制遍历深度为1:
// 添加深度参数1,仅遍历当前目录的直接子项 try (Stream<Path> paths = Files.walk(Paths.get(String.valueOf(files)), 1)){ paths.filter(Files::isRegularFile) .filter(DistributionFolder::isTXT) .forEach(/* 原有处理逻辑不变 */); }
这样每个子目录仅处理自己直接的txt文件,主类遍历所有层级的子目录时也不会出现重复读取的问题。
额外优化建议
- 简化
isTXT方法,不用拆分路径:
private static boolean isTXT(Path path) { return "txt".equals(Files.getFileExtension(path)); }
- 线程池可复用,避免每次调用方法都新建,减少资源浪费。
内容的提问来源于stack exchange,提问作者ilie alexandru
相关产品推荐
相关产品推荐

