C#中AsSpan().Fill()在不同类型数组的性能差异及原因探究
关于Span.Fill()在不同类型大数组中性能差异的解释
我发现一个特殊现象:使用AsSpan().Fill()操作时,相同字节大小的byte[]数组执行速度是Int32[]或float[]数组的两倍,但该差异仅在大尺寸数组中显现,小尺寸数组的操作速度基本一致。
复现代码
internal unsafe class Program { static byte[]? ByteFrame; static Int32[]? Int32Frame; static float[]? FloatFrame; static int[]? ResetCacheArray; static void Main(string[] args) { // size vars int Width = 1500; int Height = 1500; // Init frames ByteFrame = new byte[Width * Height * 4]; ByteFrame.AsSpan().Fill(0); Int32Frame = new Int32[Width * Height]; Int32Frame.AsSpan().Fill(0); FloatFrame = new float[Width * Height]; FloatFrame.AsSpan().Fill(1); ResetCacheArray = new int[10000 * 10000]; ResetCacheArray.AsSpan().Fill(1); // warmup jitter for(int i = 0; i < 200; i++) { ClearByteFrameAsSpanFill(0); ClearInt32FrameAsSpanFill(0); ClearFloatFrameAsSpanFill(0f); ClearCache(); } Console.WriteLine(Environment.Is64BitProcess); int TestIterations; double nanoseconds; double MsDuration; double MB = 0; double MBSec; double GBSec; TestIterations = 1; nanoseconds = 1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency; for (int i = 0; i < TestIterations; i++) { MB = ClearByteFrameAsSpanFill(0); } MsDuration = (((1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency) - nanoseconds) / TestIterations) / 1000000; MBSec = (MB / MsDuration) * 1000; GBSec = MBSec / 1000; Console.WriteLine("ClearByteFrameAsSpanFill: MS:" + MsDuration + " GB/s:" + (int)GBSec + " MB/s:" + (int)MBSec); ClearCache(); TestIterations = 1; nanoseconds = 1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency; for (int i = 0; i < TestIterations; i++) { MB = ClearInt32FrameAsSpanFill(1); } MsDuration = (((1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency) - nanoseconds) / TestIterations) / 1000000; MBSec = (MB / MsDuration) * 1000; GBSec = MBSec / 1000; Console.WriteLine("ClearInt32FrameAsSpanFill: MS:" + MsDuration + " GB/s:" + (int)GBSec + " MB/s:" + (int)MBSec); ClearCache(); TestIterations = 1; nanoseconds = 1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency; for (int i = 0; i < TestIterations; i++) { MB = ClearFloatFrameAsSpanFill(1f); } MsDuration = (((1_000_000_000.0 * Stopwatch.GetTimestamp() / Stopwatch.Frequency) - nanoseconds) / TestIterations) / 1000000; MBSec = (MB / MsDuration) * 1000; GBSec = MBSec / 1000; Console.WriteLine("ClearFloatFrameAsSpanFill: MS:" + MsDuration + " GB/s:" + (int)GBSec + " MB/s:" + (int)MBSec); ClearCache(); Console.ReadLine(); } static double ClearByteFrameAsSpanFill(byte clearValue) { ByteFrame.AsSpan().Fill(clearValue); return ByteFrame.Length / 1000000; } static double ClearInt32FrameAsSpanFill(Int32 clearValue) { Int32Frame.AsSpan().Fill(clearValue); return (Int32Frame.Length * 4) / 1000000; } static double ClearFloatFrameAsSpanFill(float clearValue) { FloatFrame.AsSpan().Fill(clearValue); return (FloatFrame.Length * 4) / 1000000; } static void ClearCache() { int sum = 0; for (int i = 0; i < ResetCacheArray.Length; i++) { sum += ResetCacheArray[i]; } } }
测试结果
小尺寸数组(1500×1500)
此时三种数组Fill操作速度相近:
ClearByteFrameAsSpanFill: MS:0,4913 GB/s:18 MB/s:18318 ClearInt32FrameAsSpanFill: MS:0,4851 GB/s:18 MB/s:18552 ClearFloatFrameAsSpanFill: MS:0,458 GB/s:19 MB/s:19650
大尺寸数组(4500×4500)
此时byte[]的Fill速度明显快于Int32[]和float[]:
ClearByteFrameAsSpanFill: MS:3,4015 GB/s:23 MB/s:23813 ClearInt32FrameAsSpanFill: MS:7,635 GB/s:10 MB/s:10609 ClearFloatFrameAsSpanFill: MS:7,4429 GB/s:10 MB/s:10882
原因解释
核心差异来自CPU的缓存行为、写合并机制和SIMD指令的利用率:
写合并(Write Combining)效率差异
CPU的写合并缓冲区(WC Buffer)允许将多个小内存写操作合并为单个总线事务,减少总线交互次数。byte类型的Fill操作可以更高效地填满WC缓冲区:单个byte值的连续写入能让CPU以最优方式合并写请求,而int/float的4字节写操作虽然单元素数据量更大,但在合并时的总线事务效率低于byte的批量写。SIMD指令的处理能力
.NET的Span.Fill()内部会用SIMD指令优化:- 针对
byte,可以用AVX2指令一次性将单个byte值广播到256位寄存器(覆盖32个byte),一次操作写入32个元素; - 针对
int/float,同样的256位寄存器只能覆盖8个元素。虽然单次操作的数据总量相同(32×1=8×4字节),但大数组场景下,内存系统对连续byte写的缓存行回写管理更高效,减少了额外的总线开销。
- 针对
缓存容量的影响
- 小尺寸数组能完全放入CPU的L3缓存,此时所有操作都在高速缓存中完成,内存访问的开销被掩盖,性能差异不明显;
- 大尺寸数组超出L3缓存容量,需要频繁与主内存交互,此时写合并和SIMD的效率差异被放大,最终显现出两倍的性能差距。
内容的提问来源于stack exchange,提问作者Lasse
相关产品推荐
相关产品推荐

