C# .NET中按文件列表分配线程数的多线程Google云上传方案
Absolutely, this is totally achievable—and it’s a smart way to optimize upload throughput by matching thread counts to file size profiles. Let’s break down the optimal implementation step by step, like how I’d approach it on a production project.
Step 1: Sort and Split Files into Capacity-Balanced Lists
First, you need to sort your files by size (descending, so big files come first) then split them into three groups where each group’s total size is roughly equal. Here’s the breakdown:
- Calculate total capacity: Sum up all file sizes to get the target per-group size (
totalSize / 3). - Iterate and accumulate: Start adding files to the first group until its total hits the target, then move to the second group, and dump the rest into the third.
- Handle edge cases: If a single file is larger than the target group size, you’ll have to leave it alone in a group (you can’t split a file, after all)—just adjust the targets for the remaining groups accordingly.
Here’s a quick code snippet for this logic:
var files = new List<FileInfo>(/* your source file collection */); files.Sort((a, b) => b.Length.CompareTo(a.Length)); // Sort descending by size long totalSize = files.Sum(f => f.Length); long targetGroupSize = totalSize / 3; var group1 = new List<FileInfo>(); var group2 = new List<FileInfo>(); var group3 = new List<FileInfo>(); long currentAccumulatedSize = 0; foreach (var file in files) { if (currentAccumulatedSize + file.Length <= targetGroupSize) { group1.Add(file); currentAccumulatedSize += file.Length; } else if (currentAccumulatedSize + file.Length <= 2 * targetGroupSize) { group2.Add(file); currentAccumulatedSize += file.Length; } else { group3.Add(file); } } // Optional tweak: If the last group is way off balance, shift a large file from group3 to group2 if (group3.Sum(f => f.Length) > targetGroupSize * 1.1 && group2.Count > 0) { var largestInGroup2 = group2.OrderByDescending(f => f.Length).First(); group3.Add(largestInGroup2); group2.Remove(largestInGroup2); }
Step 2: Parallel Processing with Per-Group Thread Limits
For fine-grained control, run each group’s uploads in parallel with its own thread limit, then wait for all groups to finish. The cleanest approach uses Parallel.ForEach with MaxDegreeOfParallelism for each group, wrapped in Task.WhenAll to run all groups concurrently.
Key Best Practices:
- Reuse Google Cloud clients: Don’t create a new
StorageClientper upload—it’s thread-safe, so reuse a single instance to avoid overhead. - Error isolation: Add try/catch blocks inside the upload loop to handle individual file failures without crashing the entire batch.
- Cancellation support: Use a
CancellationTokento safely stop all uploads if needed (e.g., user hits a cancel button).
Here’s the core upload logic:
// Initialize your Google Cloud Storage client (reuse this across all uploads!) var storageClient = StorageClient.Create(); var bucketName = "your-target-bucket-name"; var cancellationToken = new CancellationTokenSource().Token; // Define the async upload function for a single file async Task UploadSingleFile(FileInfo file) { try { Console.WriteLine($"Starting upload: {file.Name}"); using var fileStream = file.OpenRead(); await storageClient.UploadObjectAsync( bucketName, file.Name, null, fileStream, cancellationToken: cancellationToken); Console.WriteLine($"Completed upload: {file.Name}"); } catch (Exception ex) { Console.WriteLine($"Failed to upload {file.Name}: {ex.Message}"); // Add retry logic or error logging here if needed } } // Run each group with its dedicated thread limit var group1Task = Task.Run(() => { Parallel.ForEach(group1, new ParallelOptions { MaxDegreeOfParallelism = 5, CancellationToken = cancellationToken }, file => UploadSingleFile(file).Wait()); // .Wait() is safe here since we're in a dedicated Task }); var group2Task = Task.Run(() => { Parallel.ForEach(group2, new ParallelOptions { MaxDegreeOfParallelism = 3, CancellationToken = cancellationToken }, file => UploadSingleFile(file).Wait()); }); var group3Task = Task.Run(() => { Parallel.ForEach(group3, new ParallelOptions { MaxDegreeOfParallelism = 2, CancellationToken = cancellationToken }, file => UploadSingleFile(file).Wait()); }); // Wait for all groups to finish processing await Task.WhenAll(group1Task, group2Task, group3Task);
Alternative: Fully Async Parallelism (No Blocking)
If you prefer to avoid .Wait() and stick to pure async code, use a SemaphoreSlim to enforce parallelism limits per group:
async Task ProcessGroupAsync(List<FileInfo> group, int maxParallelism) { var semaphore = new SemaphoreSlim(maxParallelism); var uploadTasks = new List<Task>(); foreach (var file in group) { await semaphore.WaitAsync(cancellationToken); uploadTasks.Add(Task.Run(async () => { try { await UploadSingleFile(file); } finally { semaphore.Release(); } }, cancellationToken)); } await Task.WhenAll(uploadTasks); } // Run all groups concurrently await Task.WhenAll( ProcessGroupAsync(group1, 5), ProcessGroupAsync(group2, 3), ProcessGroupAsync(group3, 2));
This is more idiomatic for async-first C# code and avoids blocking threads unnecessarily.
Step 3: Final Optimizations
- Exponential backoff: Add retry logic for transient Google Cloud API errors (the built-in client has retry support, but you can customize it via
UploadObjectOptions). - Progress tracking: Use
IProgress<long>to report upload byte counts to your UI or logging system. - Throttle if needed: If you hit API rate limits, adjust the parallelism counts or add small delays between uploads.
内容的提问来源于stack exchange,提问作者LJLRD

