.NET 8中ConcurrentDictionary比Dictionary基准测试更快的原因解析
为什么ConcurrentDictionary.TryGetValue性能优于Dictionary.TryGetValue?
我编写了一个简单基准测试,对比Dictionary<string, int>与ConcurrentDictionary<string, int>的读取性能,测试代码如下:
[MemoryDiagnoser] public class ExtremelySimpleDictionaryBenchmark { private readonly List<string> _keys = new(); private readonly Dictionary<string, int> _dict = new(); private readonly ConcurrentDictionary<string, int> _concurrentDict = new(); private const int DictionarySize = 300; private const string Alphabet = "abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ"; private static string Generate(int length, Random random) { var result = new char[length]; for (var i = 0; i < length; i++) result[i] = Alphabet[random.Next(Alphabet.Length)]; return new string(result); } [GlobalSetup] public void Setup() { var rnd = new Random(); for (var i = 0; i < DictionarySize; i++) { var key = Generate(6, rnd); var count = rnd.Next(1000); _dict[key] = count; _concurrentDict[key] = count; _keys.Add(key); } } [Benchmark(Baseline = true)] public int Dictionary() { // variable res is used to avoid dead code elimination var res = 0; for (var index = 0; index < _keys.Count; index++) if(_dict.TryGetValue(_keys[index], out var value)) res += value; return res; } [Benchmark] public int ConcurrentDictionary() { // variable res is used to avoid dead code elimination var res = 0; for (var index = 0; index < _keys.Count; index++) if(_concurrentDict.TryGetValue(_keys[index], out var value)) res += value; return res; } }
测试环境为.NET Core 8、Windows 11、Intel Core i7-14650HX处理器,结果显示ConcurrentDictionary.TryGetValue性能反而优于Dictionary.TryGetValue,具体数据如下:
| Method | Mean | Error | StdDev | Ratio | Allocated | Alloc Ratio |
|---|---|---|---|---|---|---|
| Dictionary | 2.055 us | 0.0230 us | 0.0215 us | 1.00 | - | NA |
| ConcurrentDictionary | 1.522 us | 0.0172 us | 0.0152 us | 0.74 | - | NA |
反常结果的解释
- 无竞争场景的结构优化:ConcurrentDictionary在无竞争读取时采用无锁逻辑,.NET 8中对其读取路径做了针对性优化,比如减少内存屏障开销、更高效的哈希桶遍历逻辑。而Dictionary的TryGetValue包含了更多针对写入场景的边界检查(比如扩容状态判断),在单线程读取时反而产生额外开销。
- JIT编译优化差异:ConcurrentDictionary的TryGetValue方法代码路径更简洁,JIT编译器能更彻底地进行内联和优化;而Dictionary的内部逻辑相对复杂,某些分支可能无法被JIT完全优化,导致执行效率略低。
- 测试数据规模的影响:本次测试仅使用300个元素,ConcurrentDictionary的分段哈希结构在小数据量下的缓存命中率更高,而Dictionary的冲突链表处理在该规模下的优势未体现,反而因结构特性带来微小的性能损耗。
- 哈希算法适配性:测试中生成的随机字符串键的哈希分布可能恰好更适配ConcurrentDictionary的哈希计算方式,减少了哈希冲突的概率,进而提升了读取速度。
内容的提问来源于stack exchange,提问作者Pupkin
相关产品推荐
相关产品推荐

