为何C#中Parallel For循环比单线程循环运行更慢?
Great question! The reason your Parallel.For version underperforms compared to the simple single-threaded loop boils down to three key, relatable factors:
1. Trivial Work Per Iteration
Your single iteration does nothing but sum++—this is an extremely fast operation, barely taking a single CPU cycle. In contrast, Parallel.For has significant overhead: it needs to split the total iterations into chunks, manage thread pool threads, handle task scheduling, and coordinate between threads. All this overhead far outweighs the tiny bit of work each iteration does. The single-threaded loop skips all these costs entirely, so it's naturally faster.
2. Cache Contention (False Sharing)
The sum variable is shared across all threads in the parallel version. Every time a thread increments sum, it has to update the value in its CPU cache. But thanks to CPU cache coherence protocols (like MESI), other cores' caches that hold a copy of sum have to be invalidated and updated. This constant synchronization creates massive overhead—far more than the actual increment operation. In the single-threaded loop, sum stays in the CPU's L1 cache the whole time, with zero synchronization cost.
3. Thread Pool Scheduling Overhead
Parallel.For relies on the .NET thread pool to assign threads. Spinning up threads (or even reusing existing ones) involves context switching, which has its own cost. When each task's execution time is way smaller than the time spent scheduling it, parallelization doesn't help—it hurts.
How to Make Parallel.For Faster
If you want to see parallelization pay off, try these fixes:
- Increase Work Per Iteration: Make each iteration do meaningful work (e.g., complex calculations, processing data chunks) instead of a trivial increment. The more work each iteration does, the more the parallelization overhead gets amortized.
- Avoid Shared State: Use local accumulators to eliminate cache contention. Here's an optimized version of your parallel code:
static void Main(string[] args) { var iterations = 1000000000; var stp = new Stopwatch(); // Single-threaded baseline stp.Start(); var singleSum = 0; for (int i = 0; i < iterations; i++) { singleSum++; } stp.Stop(); Console.WriteLine("Single Thread"); Console.WriteLine($"Sum: {singleSum}"); Console.WriteLine($"Time Taken (ms): {stp.ElapsedMilliseconds}"); // Optimized Parallel.For with local accumulators stp.Restart(); long parallelSum = 0; Parallel.For( 0, iterations, () => 0L, // Initialize local accumulator per thread (i, state, localSum) => localSum + 1, // Increment local sum localSum => Interlocked.Add(ref parallelSum, localSum) // Merge local sums safely ); stp.Stop(); Console.WriteLine("\nOptimized Parallel For"); Console.WriteLine($"Sum: {parallelSum}"); Console.WriteLine($"Time Taken (ms): {stp.ElapsedMilliseconds}"); }
This version lets each thread accumulate its own local sum, so there's no constant cache synchronization. We only merge the results once per thread, which minimizes the overhead of shared state.
内容的提问来源于stack exchange,提问作者dipen

