为何C#中结构体复制速度不慢于原地修改数据?
为何ref struct原地Add方法性能略逊于+运算符?
我定义了一个表示三维点的C# ref struct:
public ref struct Point { public double X, Y, Z; // 原地加法 public void Add(Point other) { X += other.X; Y += other.Y; Z += other.Z; } // +运算符重载 public static Point operator +(Point a, Point b) { return new Point { X = a.X + b.X, Y = a.Y + b.Y, Z = a.Z + b.Z }; } }
该结构体包含两种加法实现:Add方法原地修改实例,+运算符返回新实例。我原本认为原地修改的Add方法因无需创建新实例、仅一次参数复制,性能会优于或等于+运算符,但通过BenchmarkDotNet测试发现,Add方法的平均速度反而略慢。
测试方法
AddInPlace(原地加法)
public Point AddInPlace() { Point result = new Point(); foreach ((double, double, double) coordinate in _coordinates) { Point point = new Point { X = coordinate.Item1, Y = coordinate.Item2, Z = coordinate.Item3 }; result.Add(point); // 原地修改 } return result; }
AddOperator(+运算符)
public Point AddOperator() { Point result = new Point(); foreach ((double, double, double) coordinate in _coordinates) { Point point = new Point { X = coordinate.Item1, Y = coordinate.Item2, Z = coordinate.Item3 }; result = point + result; // 使用+运算符 } return result; }
测试结果(.NET 6.0)
| Method | Job | Runtime | N | Mean | Error | StdDev |
|---|---|---|---|---|---|---|
| AddInPlace | .NET 6.0 | .NET 6.0 | 100 | 91.00 ns | 0.575 ns | 0.537 ns |
| AddOperator | .NET 6.0 | .NET 6.0 | 100 | 89.92 ns | 0.633 ns | 0.592 ns |
| AddInPlace | .NET 6.0 | .NET 6.0 | 1000 | 927.93 ns | 4.706 ns | 4.172 ns |
| AddOperator | .NET 6.0 | .NET 6.0 | 1000 | 926.57 ns | 2.606 ns | 2.310 ns |
| AddInPlace | .NET 6.0 | .NET 6.0 | 10000 | 9,379.43 ns | 106.127 ns | 94.079 ns |
| AddOperator | .NET 6.0 | .NET 6.0 | 10000 | 9,303.92 ns | 18.739 ns | 16.611 ns |
测试环境:
BenchmarkDotNet v0.13.12, Windows 11 Intel Xeon W-2235 CPU 3.80GHz, 1 CPU, 12 logical and 6 physical cores
原因分析
两者性能差异极小(处于误差边缘),但AddInPlace稍慢的核心原因在于内存操作模式的差异:
- Add方法的读-改-写(RMW)操作:Add方法执行
X += other.X这类逻辑时,是三次独立的「读取字段值→计算→写回字段」流程。这类操作需要CPU先读取内存值、计算后再写回,流水线效率较低,还可能引发缓存写回开销。 - +运算符的直接赋值:
+运算符被JIT完全内联后,会先计算三个维度的和,再一次性将结果赋值给result的三个字段。这种「先计算所有值→批量赋值」的模式避免了RMW操作的等待周期,内存访问更连续,CPU流水线可更好地并行处理加法与赋值。 - JIT优化的细微差异:对于ref struct,静态运算符的参数传递与返回值优化(RVO)能被JIT更彻底处理,编译器可直接在寄存器中完成加法计算,减少栈上临时复制;而实例方法的
this指针传递与字段修改,会引入微小的额外栈操作开销。
内容的提问来源于stack exchange,提问作者Lorenzo Celli
相关产品推荐
相关产品推荐

