整数向量化除法精度是否依赖CPU?.NET向量化计算异常咨询
16位ARGB整数预乘向量化操作的精度与运行时问题
我正在对16位整数ARGB通道的64位颜色执行预乘向量化操作。由于缺乏加速整数除法支持,需要将值转为float并显式使用SSE2/SSE4.1 intrinsics以获取最佳性能,但同时保留了通用版本作为降级方案(尽管当前性能低于常规操作,但可兼容未来优化)。
最小复现:精度偏差问题
// Test color with 50% alpha (ushort A, ushort R, ushort G, ushort B) c = (0x8000, 0xFFFF, 0xFFFF, 0xFFFF); // Minimal version of the fallback logic if HW intrinsics cannot be used: Vector128<uint> v = Vector128.Create(c.R, c.G, c.B, 0u); v = v * c.A / Vector128.Create(0xFFFFu); var cPre = (c.A, (ushort)v[0], (ushort)v[1], (ushort)v[2]); // Original color: Console.WriteLine(c); // prints (32768, 65535, 65535, 65535) // Expected premultiplied color: (32768, 32768, 32768, 32768) Console.WriteLine(cPre); // prints (32768, 32769, 32769, 32769)
运行结果不符合预期,且SharpLab中结果正确,但.NET Fiddle可复现该问题。疑问:这是特定平台的预期表现还是需提交运行时bug?
扩展测试:结果完全错误
using System; using System.Numerics; using System.Runtime.Intrinsics; (ushort A, ushort R, ushort G, ushort B) c = (32768, 65535, 32768, 16384); Vector128<uint> v1 = Vector128.Create(c.R, c.G, c.B, 0u); v1 = v1 * c.A / Vector128.Create(0xFFFFu); // prints <32769, 49152, 57344, 0> instead of <32768, 16384, 8192, 0> Console.WriteLine(v1); // Also for the older Vector<T> Span<uint> span = stackalloc uint[Vector<uint>.Count]; span[0] = c.R; span[1] = c.G; span[2] = c.B; Vector<uint> v2 = new Vector<uint>(span) * c.A / new Vector<uint>(0xFFFF); // prints <32769, 49152, 57344, 0, 0, 0, 0, 0> on my machine Console.WriteLine(v2);
问题根源:ushort值的错误处理
最终发现问题出在乘法阶段:将* c.A替换为常量表达式* 32768后结果正确,ushort值未被正确提取/掩码,甚至Vector.Create也受影响:
(ushort A, ushort R, ushort G, ushort B) c = (32768, 65535, 32768, 16384); Console.WriteLine(Vector128.Create((int)c.A)); // -32768 Console.WriteLine(Vector128.Create((int)32768)); // 32768 Console.WriteLine(Vector128.Create((int)c.A, (int)c.A, (int)c.A, (int)c.A)); // 32768
已在.NET运行时仓库提交issue,同时咨询:整数向量化除法的精度/准确度是否依赖CPU?
内容的提问来源于stack exchange,提问作者György Kőszeg
相关产品推荐
相关产品推荐

