Parallel文件夹大小扫描WaitAll()挂起,Main无法异步如何解决?
Hey Justin, let's tackle this problem head-on. You're right that skipping await on async calls from a synchronous Main can lead to deadlocks (or just hanging), and since you can't make Main async, we need workarounds that play nice with synchronous entry points. Plus, we'll optimize your folder size checking to be as fast as possible—including supporting out-of-order CSV writes if that helps speed things up.
First, let's ground this in your code context. I'll assume your original synchronous code and problematic async attempt look something like these (feel free to adjust if I'm off-base):
Your Original Synchronous Program
static void Main(string[] args) { var targetDirs = new List<string> { @"\\server1\shared\docs", @"\\server2\shared\media", @"\\server3\shared\backup" }; var sizeResults = new List<DirSizeResult>(); foreach (var dir in targetDirs) { long totalSize = CalculateDirSizeSync(dir); sizeResults.Add(new DirSizeResult { Path = dir, SizeInBytes = totalSize }); } WriteResultsToCsv(sizeResults, @"folder_sizes.csv"); } static long CalculateDirSizeSync(string path) { long total = 0; try { foreach (var file in Directory.GetFiles(path, "*", SearchOption.AllDirectories)) { total += new FileInfo(file).Length; } } catch (Exception ex) { Console.WriteLine($"Failed to scan {path}: {ex.Message}"); } return total; } public class DirSizeResult { public string Path { get; set; } public long SizeInBytes { get; set; } }
Your Problematic Async Attempt (Causing Hangs/Deadlocks)
static void Main(string[] args) { var targetDirs = new List<string> { /* ... same paths ... */ }; var asyncTasks = new List<Task<DirSizeResult>>(); foreach (var dir in targetDirs) { asyncTasks.Add(CalculateDirSizeAsync(dir)); } // Using .Result or .Wait() here can trigger deadlocks, especially if async code captures sync context var results = asyncTasks.Select(t => t.Result).ToList(); WriteResultsToCsv(results, @"folder_sizes.csv"); } static async Task<DirSizeResult> CalculateDirSizeAsync(string path) { long total = 0; try { foreach (var file in Directory.EnumerateFiles(path, "*", SearchOption.AllDirectories)) { // Trying to use async IO, but still hitting issues using var stream = new FileStream(file, FileMode.Open, FileAccess.Read, FileShare.Read, 4096, useAsync: true); total += stream.Length; } } catch (Exception ex) { Console.WriteLine($"Failed to scan {path}: {ex.Message}"); } return new DirSizeResult { Path = path, SizeInBytes = total }; }
Step 1: Fix the Deadlock in Non-Async Main
The core issue here is that when you call .Result or .Wait() on an async task in a synchronous context, you can block the thread that the async code needs to resume on (especially if the async code captures a synchronization context, though console apps have a simpler context than UI apps). Here are two reliable fixes:
Option 1: Use GetAwaiter().GetResult() Instead of .Result
This avoids wrapping exceptions in AggregateException and is safer for console apps to prevent deadlocks:
static void Main(string[] args) { var targetDirs = new List<string> { /* ... paths ... */ }; var asyncTasks = targetDirs.Select(CalculateDirSizeAsync).ToList(); // Wait for all tasks to complete and get results safely var allResults = Task.WhenAll(asyncTasks).GetAwaiter().GetResult(); WriteResultsToCsv(allResults.ToList(), @"folder_sizes.csv"); }
Option 2: Wrap Async Work in Task.Run
If you're still hitting deadlocks (unlikely in console apps, but possible if your async code uses context-bound operations), offload the async work to a thread pool thread:
static void Main(string[] args) { var targetDirs = new List<string> { /* ... paths ... */ }; var poolTasks = targetDirs.Select(dir => Task.Run(() => CalculateDirSizeAsync(dir))).ToList(); // Wait for all thread pool tasks to finish Task.WaitAll(poolTasks); var allResults = poolTasks.Select(t => t.Result).ToList(); WriteResultsToCsv(allResults.ToList(), @"folder_sizes.csv"); }
Step 2: Optimize for Maximum Speed (Out-of-Order CSV Writes Allowed)
Since you're okay with out-of-order CSV writes, we can skip waiting for all tasks to finish before writing results. This reduces memory usage and lets you see progress as tasks complete. Just make sure to handle thread safety for CSV writes!
Optimized Async Code with Real-Time CSV Writes
// Lock to ensure safe concurrent writes to the CSV file private static readonly object _csvWriteLock = new object(); // Limit parallelism to avoid overwhelming network/shared storage private static readonly SemaphoreSlim _concurrencyLimiter = new SemaphoreSlim(8); static void Main(string[] args) { var targetDirs = new List<string> { /* ... paths ... */ }; // Write CSV header first lock (_csvWriteLock) { using var writer = new StreamWriter(@"folder_sizes.csv", append: false); writer.WriteLine("DirectoryPath,SizeInBytes,Status"); } var asyncTasks = targetDirs.Select(async dir => { await _concurrencyLimiter.WaitAsync(); DirSizeResult result = null; string status = "Success"; try { result = await CalculateDirSizeAsync(dir); } catch (Exception ex) { status = $"Failed: {ex.Message}"; Console.WriteLine($"Error scanning {dir}: {ex.Message}"); } finally { _concurrencyLimiter.Release(); } // Write result immediately after task completes lock (_csvWriteLock) { using var writer = new StreamWriter(@"folder_sizes.csv", append: true); writer.WriteLine($"{EscapeCsvValue(dir)},{result?.SizeInBytes ?? 0},{EscapeCsvValue(status)}"); } }).ToList(); // Wait for all tasks to finish Task.WhenAll(asyncTasks).GetAwaiter().GetResult(); Console.WriteLine("All directory scans completed!"); } static async Task<DirSizeResult> CalculateDirSizeAsync(string path) { long totalSize = 0; // Use async file enumeration (available in .NET Core 3.0+) for better memory efficiency await foreach (var filePath in Directory.EnumerateFilesAsync(path, "*", SearchOption.AllDirectories)) { // Use async file info to get size (available in .NET 5+) var fileInfo = new FileInfo(filePath); totalSize += await fileInfo.LengthAsync; } return new DirSizeResult { Path = path, SizeInBytes = totalSize }; } // Helper to escape CSV values with commas/quotes static string EscapeCsvValue(string value) { if (value.Contains(',') || value.Contains('"')) { return $"\"{value.Replace("\"", "\"\"")}\""; } return value; }
Key Optimizations Here:
- Concurrency Limiting: The
SemaphoreSlimprevents spamming too many concurrent network/IO requests, which can slow down both your app and the shared servers. Adjust the number (8 in the example) based on your network/storage capacity. - Async File Enumeration:
Directory.EnumerateFilesAsyncstreams file paths instead of loading all into memory at once—critical for large directories. - Real-Time CSV Writes: No need to store all results in memory; write as soon as a scan finishes.
- Thread-Safe Writes: The
lockensures multiple tasks don't corrupt the CSV file. - Proper Error Handling: Failed scans are logged and marked in the CSV instead of crashing the whole app.
Bonus: Even Faster Scanning
If you want to push performance further:
- Batch File Size Reads: For directories with thousands of files, batch async reads to reduce overhead.
- Use
FileInfo.LengthAsync: As shown, this is more efficient than opening a stream just to get the file size. - Avoid Redundant Error Logging: If you're writing errors to CSV, you can skip console logging (or vice versa) to save a tiny bit of time.
内容的提问来源于stack exchange,提问作者Justin Rice

