Scala中如何基于允许文件列表最高效处理文件夹内的目标文件?
方案建议
核心优化点
首先把你的允许文件列表从List改成Set,文件名存在性判断的时间复杂度会从O(n)降到O(1),文件数量多的时候性能提升非常明显。
方案1:递归实现(支持任意深度目录遍历)
不需要写双层递归,单层递归就可以覆盖所有层级的目录遍历,同时把文件处理函数作为参数传入,完全解耦遍历逻辑和业务处理逻辑:
import java.io.File /** * 遍历目录下所有符合条件的文件并执行处理逻辑 * @param rootDir 遍历根目录 * @param allowedFileNames 允许处理的文件名集合 * @param process 自定义文件处理函数 */ def processAllowedFiles(rootDir: File, allowedFileNames: Set[String], process: File => Unit): Unit = { if (rootDir.exists && rootDir.isDirectory) { val entries = rootDir.listFiles() // 兼容无权限访问、目录被删除等特殊场景,避免空指针 if (entries != null) { entries.foreach { entry => if (entry.isFile) { if (allowedFileNames.contains(entry.getName)) { process(entry) } } else if (entry.isDirectory) { // 子目录递归遍历 processAllowedFiles(entry, allowedFileNames, process) } } } } }
使用示例
// 你自己的业务处理逻辑 def customProcess(file: File): Unit = { println(s"处理文件:${file.getAbsolutePath}") // 此处补充你的业务处理代码 } // 调用入口 val rootDirectory = new File("/your/target/root/path") val allowedFiles = Set("File1.json", "File2.json", "File3.json") processAllowedFiles(rootDirectory, allowedFiles, customProcess)
方案2:双层遍历(仅适用于固定两层目录的场景)
如果你确定只需要遍历根目录下的一级子目录,不需要处理更深层级的目录,直接用集合操作实现更简洁,性能和递归基本一致:
import java.io.File val rootDirectory = new File("/your/target/root/path") val allowedFiles = Set("File1.json", "File2.json", "File3.json") def customProcess(file: File): Unit = { // 你的业务处理逻辑 } if (rootDirectory.exists && rootDirectory.isDirectory) { rootDirectory.listFiles() .filter(_.isDirectory) // 筛选根目录下的一级子文件夹 .flatMap(dir => Option(dir.listFiles()).getOrElse(Array.empty)) // 取出子文件夹下所有内容,兼容空指针场景 .filter(file => file.isFile && allowedFiles.contains(file.getName)) // 筛选符合要求的文件 .foreach(customProcess) // 批量处理 }
内容的提问来源于stack exchange,提问作者Kylo
相关产品推荐
相关产品推荐

