递归遍历父目录收集含指定后缀文件目录的优化方案问询
Great question! Your initial approach gets the job done, but it’s inefficient because it first fetches every matching file (even duplicates in the same directory) and then deduplicates the parent directories afterward. Let’s look at a few cleaner, more performant alternatives:
1. Check Directories First (Most Efficient)
Instead of grabbing all files first, iterate through each directory and check if it contains any matching files. This avoids redundant work with multiple files in the same directory, and uses lazy enumeration to keep memory usage low:
var testDirectories = new List<DirectoryInfo>(); // Enumerate all subdirectories recursively (lazy-loaded) foreach (var dirPath in Directory.EnumerateDirectories(parentDirectory, "*", SearchOption.AllDirectories)) { // Check if this directory has any matching files (stops at the first match) if (Directory.EnumerateFiles(dirPath, "*spec.js").Any()) { testDirectories.Add(new DirectoryInfo(dirPath)); } }
Why this is better:
- Uses
EnumerateDirectoriesandEnumerateFilesinstead ofGetDirectories/GetFiles—these methods don’t load all paths into memory upfront, which is a big win for large directory structures. - Only processes each directory once, even if it has dozens of matching files.
Any()short-circuits as soon as it finds one matching file, so you don’t waste time checking all files in the directory.
2. LINQ Version (Clean & Concise)
If you prefer a more compact syntax, you can wrap the above logic in a LINQ chain:
var testDirectories = Directory.EnumerateDirectories(parentDirectory, "*", SearchOption.AllDirectories) .Where(dirPath => Directory.EnumerateFiles(dirPath, "*spec.js").Any()) .Select(dirPath => new DirectoryInfo(dirPath)) .ToList();
This does exactly the same thing as the first approach, just with a more declarative style.
3. Custom Recursion (For Advanced Control)
If you need to add custom rules (like skipping certain directories, e.g., node_modules or bin), a manual recursive method gives you full control:
private static List<DirectoryInfo> GetTestDirectories(string parentDir, string filePattern) { var matchingDirs = new List<DirectoryInfo>(); // Check if the current directory has matching files if (Directory.EnumerateFiles(parentDir, filePattern).Any()) { matchingDirs.Add(new DirectoryInfo(parentDir)); } // Recurse into subdirectories, with optional skip logic foreach (var subDirPath in Directory.EnumerateDirectories(parentDir)) { // Example: Skip node_modules directories if (Path.GetFileName(subDirPath).Equals("node_modules", StringComparison.OrdinalIgnoreCase)) continue; matchingDirs.AddRange(GetTestDirectories(subDirPath, filePattern)); } return matchingDirs; } // Usage: var testDirectories = GetTestDirectories(parentDirectory, "*spec.js");
This is perfect if you need to tailor the traversal to your specific project structure.
Quick Comparison to Your Original Code
Your initial approach:
- Fetches every
*.spec.jsfile in the entire tree - Adds their parent directory to a list (duplicating entries for directories with multiple matching files)
- Deduplicates the list with
Distinct()
This does unnecessary work—if a directory has 50 spec files, you’re adding the same directory 50 times before removing duplicates. The alternatives above avoid this entirely by checking each directory once.
内容的提问来源于stack exchange,提问作者Oblomov

