为何采用CPU缓存优化的ECS基准测试未达预期性能提升?
结构体/ECS实现性能未达预期的疑问
在C#中尝试用结构体或数组实现Entity Component System(ECS),但性能相比类和对象并无明显优势。尽管采用了CPU缓存和数据局部性等优化技术,BenchmarkDotNet的测试结果却未呈现预期的性能提升。疑惑是自身实现存在问题,还是该设计在当前软硬件环境下的影响已减弱?
基准测试结果
BenchmarkDotNet=v0.13.4, OS=Windows 11 (10.0.22621.963) AMD Ryzen 5 5600X, 1 CPU, 12 logical and 6 physical cores .NET SDK=7.0.102 [Host] : .NET 7.0.2 (7.0.222.60605), X64 RyuJIT AVX2 DefaultJob : .NET 7.0.2 (7.0.222.60605), X64 RyuJIT AVX2 | Method | Mean | Error | StdDev | |---------------- |---------:|--------:|--------:| | Structs | 128.2 us | 0.83 us | 0.78 us | | Classes | 122.5 us | 0.17 us | 0.15 us | | ComponentArrays | 203.6 us | 0.53 us | 0.49 us |
测试代码
internal class Program { static void Main(string[] args) { BenchmarkRunner.Run<Benchmark>(); } } struct StructEntity { public int Age; public Vector3 Position; public float Health; } class ClassEntity { public int Age; public Vector3 Position; public float Health; } public class Benchmark { private readonly StructEntity[] _structs; private readonly ClassEntity[] _classes; private readonly int[] _ageComponents; private readonly Vector3[] _positionComponents; private readonly float[] _healthComponents; private const int size = 50000; private static Random random = new(); public Benchmark() { _structs = new StructEntity[size]; _classes = new ClassEntity[size]; _ageComponents = new int[size]; _positionComponents = new Vector3[size]; _healthComponents = new float[size]; for (var i = 0; i < _structs.Length; i++) { var age = random.Next(1, 100); var health = (float)random.NextDouble(); var position = new Vector3((float)random.NextDouble(), (float)random.NextDouble(), (float)random.NextDouble()); // structs var structEntity = new StructEntity(); structEntity.Age = age; structEntity.Health = health; structEntity.Position = position; _structs[i] = structEntity; // classes var classEntity = new ClassEntity(); _classes[i] = classEntity; classEntity.Age = age; classEntity.Health = health; classEntity.Position = position; // component arrays _healthComponents[i] = health; _ageComponents[i] = age; _positionComponents[i] = position; } } [Benchmark] public int Structs() { int count = 0; for (var i = 0; i < size; i++) { ref var structEntity = ref _structs[i]; if (structEntity.Age > 30 && structEntity.Health < 0.5) { count++; structEntity.Position = new Vector3(structEntity.Age, 101, structEntity.Age * 2); structEntity.Age *= 3; structEntity.Health *= 3; } } return count; } [Benchmark] public int Classes() { int count = 0; for (var i = 0; i < size; i++) { var classEntity = _classes[i]; if (classEntity.Age > 30 && classEntity.Health < 0.5) { count++; classEntity.Position = new Vector3(classEntity.Age, 101, classEntity.Age * 2); classEntity.Age *= 3; classEntity.Health *= 3; } } return count; } [Benchmark] public int ComponentArrays() { int count = 0; for (var i = 0; i < size; i++) { ref Vector3 position = ref _positionComponents[i]; ref int age = ref _ageComponents[i]; ref float health = ref _healthComponents[i]; if (age > 30 && health < 0.5 && position.X < position.Z) { count++; position = new Vector3(age, 101, age * 2); age *= 3; health *= 3; } } return count; } }
问题分析与解释
1. 测试场景未匹配ECS的优势场景
ECS(按组件数组存储的SOA模式)的核心优势是仅访问部分组件的场景,比如一个系统只处理Health组件时,SOA可以只加载Health数组到缓存,避免加载无关的Age、Position数据,提升缓存利用率。但你的测试中,三种实现都需要遍历所有实体并访问全部三个组件:
- 结构体数组(AOS)和类数组(引用型AOS)可以一次性加载整个实体的所有数据到缓存行,缓存利用率高。
- 组件数组(SOA)需要跨三个数组取数据,每个循环要访问三个不同的内存区域,缓存行的利用率反而下降,导致性能最差。
2. 数据规模过小,缓存优势未体现
你的测试数据量为50000个实体,每个结构体仅20字节,总数据量约1MB,完全可以放进Ryzen 5 5600X的L2缓存(每个核心512KB,总3MB)。此时缓存局部性的差异被抹平,结构体、类、SOA的缓存命中率都接近100%,自然看不到性能差距。
3. 类数组的堆分配优化
.NET的GC在分配小对象时,会在Gen0堆上连续分配,所以类数组中的对象引用虽然指向堆,但实际对象的内存地址是连续的,访问时的缓存命中率和结构体数组差距很小。加上JIT对类字段访问的优化,甚至可能出现类数组略快的情况。
改进建议
- 调整测试场景:模拟真实ECS的系统拆分,比如写一个仅统计
Health < 0.5的系统,此时SOA的Health数组会比结构体/类数组快很多,因为只需要加载单一组件数组。 - 放大数据规模:将
size改为100万甚至更大,让总数据量超过L2缓存大小,此时缓存局部性的影响会显著体现。 - 拆分系统职责:分别实现处理
Age、Health、Position的独立系统,对比SOA和AOS的性能差异。
内容的提问来源于stack exchange,提问作者Pepernoot
相关产品推荐
相关产品推荐

