Java 11下高效递归查找符合正则的文件夹及排除指定目录
Nice question! Let's break down how to optimize your existing implementation and add support for excluding specific directories, all while keeping things efficient with Java NIO.
First: Optimize Your Existing Files.find Implementation
Your initial code works, but we can tweak it to avoid unnecessary overhead and leverage NIO more effectively:
Key Optimizations:
- Skip
Path→Fileconversion: Java NIO provides all the methods we need directly onPathandBasicFileAttributes, so converting toFileadds unnecessary object creation and overhead. - Precompile your regex: Using
Pattern.compile()once instead ofString.matches()(which recompiles the regex every time) saves significant time when matching hundreds/thousands of directories. - Simplify condition checks: Combine checks for directory type, hidden status, and regex match in a cleaner, more readable way.
Optimized Code:
import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; import java.nio.file.attribute.BasicFileAttributes; import java.util.regex.Pattern; import java.util.stream.Stream; public class DirectorySearcher { public static Stream<Path> findMatchingDirs(String rootDir, String regex) { // Precompile regex for repeated use Pattern dirPattern = Pattern.compile(regex); return Files.find( Paths.get(rootDir), Integer.MAX_VALUE, (path, attrs) -> { // Skip non-directories and hidden directories first if (!attrs.isDirectory() || Files.isHidden(path)) { return false; } // Get directory name and check regex match Path dirName = path.getFileName(); return dirName != null && dirPattern.matcher(dirName.toString()).matches(); } ); } }
Adding Directory Exclusion (The Efficient Way)
The Files.find method is great for filtering, but it doesn't let you skip entire directory subtrees early. To exclude specific directories and avoid traversing their contents entirely (a big efficiency win), use Files.walkFileTree with a custom SimpleFileVisitor.
This approach lets us decide at the start of visiting a directory whether to skip its subtrees, eliminating unnecessary IO operations on excluded directories.
Code with Exclusion Support:
import java.nio.file.FileVisitResult; import java.nio.file.Files; import java.nio.file.Path; import java.nio.file.Paths; import java.nio.file.SimpleFileVisitor; import java.nio.file.attribute.BasicFileAttributes; import java.util.ArrayList; import java.util.List; import java.util.Set; import java.util.regex.Pattern; public class DirectorySearcherWithExclusion { public static List<Path> findMatchingDirs(String rootDir, String regex, Set<String> excludedDirNames) { Pattern dirPattern = Pattern.compile(regex); List<Path> matchingPaths = new ArrayList<>(); try { Files.walkFileTree(Paths.get(rootDir), new SimpleFileVisitor<>() { @Override public FileVisitResult preVisitDirectory(Path dir, BasicFileAttributes attrs) { String currentDirName = dir.getFileName().toString(); // Skip this directory and all its subtrees if it's in the exclusion list if (excludedDirNames.contains(currentDirName)) { return FileVisitResult.SKIP_SUBTREE; } // Add to results if it matches the regex and isn't hidden if (dirPattern.matcher(currentDirName).matches() && !Files.isHidden(dir)) { matchingPaths.add(dir); } // Continue traversing other subdirectories return FileVisitResult.CONTINUE; } }); } catch (Exception e) { // Handle IO exceptions (e.g., permission issues) as needed e.printStackTrace(); } return matchingPaths; } }
Why This Is More Efficient:
- Early subtree skipping: We don't waste time traversing files and subdirectories inside excluded folders—we skip them entirely at the first check.
- Minimal overhead: No unnecessary object conversions, and regex is precompiled for fast matching.
- Controlled traversal: You have full control over the visit flow, making it easy to add additional logic (like logging, permission checks) if needed.
Usage Example
public static void main(String[] args) { // Find all directories matching ".*docs.*" under "myFolder", excluding "node_modules" and ".git" Set<String> excluded = Set.of("node_modules", ".git"); List<Path> matches = DirectorySearcherWithExclusion.findMatchingDirs("myFolder", ".*docs.*", excluded); matches.forEach(System.out::println); }
内容的提问来源于stack exchange,提问作者Gregor

