You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何C#中结构体复制速度不慢于原地修改数据?

为何ref struct原地Add方法性能略逊于+运算符?

我定义了一个表示三维点的C# ref struct:

public ref struct Point
{
    public double X, Y, Z;

    // 原地加法
    public void Add(Point other)
    {
        X += other.X;
        Y += other.Y;
        Z += other.Z;
    }
    
    // +运算符重载
    public static Point operator +(Point a, Point b)
    {
        return new Point { X = a.X + b.X, Y = a.Y + b.Y, Z = a.Z + b.Z };
    }
}

该结构体包含两种加法实现:Add方法原地修改实例,+运算符返回新实例。我原本认为原地修改的Add方法因无需创建新实例、仅一次参数复制,性能会优于或等于+运算符,但通过BenchmarkDotNet测试发现,Add方法的平均速度反而略慢。

测试方法

AddInPlace(原地加法)

public Point AddInPlace()
{
    Point result = new Point();

    foreach ((double, double, double) coordinate in _coordinates)
    {
        Point point = new Point
        {
            X = coordinate.Item1, 
            Y = coordinate.Item2, 
            Z = coordinate.Item3
        };

        result.Add(point); // 原地修改
    }

    return result;
}

AddOperator(+运算符)

public Point AddOperator()
{
    Point result = new Point();

    foreach ((double, double, double) coordinate in _coordinates)
    {
        Point point = new Point
        {
            X = coordinate.Item1, 
            Y = coordinate.Item2, 
            Z = coordinate.Item3
        };

        result = point + result; // 使用+运算符
    }

    return result;
}

测试结果(.NET 6.0)

MethodJobRuntimeNMeanErrorStdDev
AddInPlace.NET 6.0.NET 6.010091.00 ns0.575 ns0.537 ns
AddOperator.NET 6.0.NET 6.010089.92 ns0.633 ns0.592 ns
AddInPlace.NET 6.0.NET 6.01000927.93 ns4.706 ns4.172 ns
AddOperator.NET 6.0.NET 6.01000926.57 ns2.606 ns2.310 ns
AddInPlace.NET 6.0.NET 6.0100009,379.43 ns106.127 ns94.079 ns
AddOperator.NET 6.0.NET 6.0100009,303.92 ns18.739 ns16.611 ns

测试环境:

BenchmarkDotNet v0.13.12, Windows 11
Intel Xeon W-2235 CPU 3.80GHz, 1 CPU, 12 logical and 6 physical cores

原因分析

两者性能差异极小(处于误差边缘),但AddInPlace稍慢的核心原因在于内存操作模式的差异:

  • Add方法的读-改-写(RMW)操作:Add方法执行X += other.X这类逻辑时,是三次独立的「读取字段值→计算→写回字段」流程。这类操作需要CPU先读取内存值、计算后再写回,流水线效率较低,还可能引发缓存写回开销。
  • +运算符的直接赋值:+运算符被JIT完全内联后,会先计算三个维度的和,再一次性将结果赋值给result的三个字段。这种「先计算所有值→批量赋值」的模式避免了RMW操作的等待周期,内存访问更连续,CPU流水线可更好地并行处理加法与赋值。
  • JIT优化的细微差异:对于ref struct,静态运算符的参数传递与返回值优化(RVO)能被JIT更彻底处理,编译器可直接在寄存器中完成加法计算,减少栈上临时复制;而实例方法的this指针传递与字段修改,会引入微小的额外栈操作开销。

内容的提问来源于stack exchange,提问作者Lorenzo Celli

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 17:42:02