并行与串行循环测试结果异常及二次执行耗时增加问题咨询
First, let's recap your test setup to make sure we're aligned: you’re comparing nested serial vs parallel loops across four scenarios—single parallel (p), single serial (s), parallel-then-serial (pfsl), and serial-then-parallel (sfpl). Your key observation is that whichever method runs second in the combined scenarios consistently takes longer—and there are concrete reasons tied to your code’s behavior and .NET runtime mechanics.
1. Accumulating State in Persistent Collections
The biggest culprit here is the unreset state stored across test runs, specifically in two critical places:
Path.Paths: A staticConcurrentDictionary<string, Path>that gets new entries added every timeFunccreates aPathinstance.Nue.pathlist: AConcurrentBag<Path>that also receives newPathinstances via thePathconstructor.
When you run the first method (serial or parallel), it generates hundreds or thousands of Path objects and adds them to these collections. By the time the second method runs:
- Traversal overhead spikes: The
pathlistin eachNueinstance is now much larger, so your innerforeach/Parallel.ForEachloops have far more items to iterate over. - ConcurrentDictionary operations slow down:
Paths.TryAddhas to check against a growing set of keys, increasing lock contention (even for concurrent collections) and adding overhead to every insertion. - Memory pressure & GC delays: The heap is cluttered with existing
PathandList<Conn>objects. Allocating new objects in the second run becomes slower due to heap fragmentation, and the garbage collector will likely trigger more frequently mid-execution, adding unexpected pause times to your total runtime.
Looking at your test numbers:
- For
pfsl(parallel-then-serial), the serial run’s time (~8k ms) is inflated compared to what it would be in isolation (your single serialsruns at 10-13k ms, but noteuntil=5means fewer loop iterations thanuntil=4—the gap would be even starker with identical loop counts). - For
sfpl(serial-then-parallel), the parallel run’s time (~6k ms) is significantly higher than the first serial run (~1.7k ms), and aligns with the pattern of second-run slowdown.
2. Thread Pool & CPU Cache Dynamics
While thread pool warming can help parallel runs on first execution, the reverse happens in combined scenarios:
- After a parallel run, the thread pool retains extra active threads. When you run the serial method next, these idle-but-warm threads can cause unnecessary context switching as the runtime manages them, adding overhead to the serial execution.
- Conversely, after a serial run, the CPU cache is filled with data from the first execution. The parallel run that follows has to process new data, leading to lower cache hit rates and slower memory access times.
3. Unreset Test Context
Your timing code doesn’t reset the application state between runs—you only restart the stopwatch. Isolated tests start with a clean heap and empty collections, so the runtime has minimal overhead to manage. Combined tests, however, run the second method on a "dirty" state with accumulated objects and potential pending GC operations, which directly skews the measured time.
How to Verify These Hypotheses
To confirm these explanations, adjust your test setup:
- Reset state between runs: Before each test (even in combined scenarios), clear
Path.Pathsand empty everyNue.pathlist(add a clear method toNueor recreate thenuelistentirely). - Force GC between runs: Add
GC.Collect(); GC.WaitForPendingFinalizers();before restarting the stopwatch for the second method. This cleans up unused objects and reduces memory pressure. - Test identical loop counts: Run combined scenarios with the same
untilvalue as your isolated tests to eliminate loop count as a variable.
内容的提问来源于stack exchange,提问作者Atomic

