MSVC中float erfcf(x)比double erfc(x)慢2倍?求复现及GCC优化建议
关于MSVC 17.1中
erfcf(x)性能异常的验证与GCC优化建议 我需要验证在MSVC 17.1标准库中观察到的一个性能异常:float erfcf(x)的速度比double erfc(x)慢2倍,确认这是普遍现象而非基准测试偏差;同时寻求GCC编译代码时的优化命令行参数建议。
我采用@Njuffa的测试框架对erfcf(x)函数开展实验,为跟踪进度添加了进度打印,发现线性遍历x的所有位模式时,打印会在特定值范围停滞,以此指导优化。测试中Intel 2024.1 ICX库的实现兼具速度与精度,但MSVC存在明显性能异常,且GCC 13.1编译的dummy循环测试速度极慢,怀疑是编译器选项设置或虚拟机环境导致,希望得到原生Linux系统的基准测试数据对比。
最小可复现代码
#include <stdio.h> #include <stdint.h> #include <string.h> #include <time.h> #include <math.h> //#define RANDOM (1) // helper routines float uint32_as_float(uint32_t a) { float r; memcpy(&r, &a, sizeof r); return r; } uint32_t float_as_uint32(float a) { uint32_t r; memcpy(&r, &a, sizeof r); return r; } bool my_isnan(float x) { // used here so that -ffast-math can't optimise it away (as it does with system isnan() // other options can be used here // return __isnan((double)x); // return __isnanf(x); uint32_t ix = float_as_uint32(x); // return !((~ix) & 0x7f800000); // Intel's choice is slightly faster return (ix & 0x7f800000) == 0x7f800000; } float erfcdf(float x) { return (float)erfc((double)x); // shim for calling double precision erfc } float dummy(float x) { return x; // to determine loop overheads } void timefun(const char* name, float (*test_fun)(float), bool verbose) { uint32_t argi, largi = 0; float arg, res, sum; time_t start0, start, end; printf("\nTiming %s\n", name); argi = 0; sum = 0.0; start0 = start = clock(); do { arg = uint32_as_float(argi); res = (*test_fun)(arg); if (!my_isnan(res)) sum += res; #ifdef RANDOM argi = (argi * 1664525 + 1013904223); // ranqd1 #else argi++; #endif if (verbose && ((argi & 0xff800000) != largi)) { end = clock(); largi = argi & 0xff800000; printf("Exp %x : %6.3f\n", argi >> 20, (float)(end - start) / CLOCKS_PER_SEC); start = clock(); } if ((argi & 0x3ffffff) == 0) printf("."); } while (argi); end = clock(); printf("\ntime taken %6.2f sum = %g\n", (float)(end - start0) / CLOCKS_PER_SEC, sum); } int main(void) { timefun("dummy", dummy, false); timefun("erfcf", erfcf, false); timefun("erfcdf", erfcdf, false); return 0; }
该代码可在GCC、Intel、MSVC编译器上直接编译,也欢迎其他编译器的基准测试结果。
基准测试数据
(线性测试遍历x的所有位模式,随机测试用ranqd1打乱分支预测,单位:秒)
| 编译器 | Dummy线性测试 | 线性erfcf | 线性erfc | 随机Dummy | 随机erfcf | 随机erfc |
|---|---|---|---|---|---|---|
| gcc 13.1 | 31 | 31.6 | 51.1 | 38.7 | 192 | 169 |
| Intel 2024.1 | 2.0 | 37.6 | 45.0 | 3.9 | 46.9 | 54.3 |
| MS 17.1 | 2.0 | 96.2 | 47.6 | 4.0 | 275 | 125 |
分析与需求
- MSVC异常分析:推测性能慢的原因是对非正规数(denorms)的多项式处理过于简单,导致速度骤降;精度测试显示,MSVC的
float erfcf(x)仅达5.7 ULP,而所有编译器的double erfc(x)转为float后精度为0.5 ULP,Intel的float erfcf(x)更是达到亚ULP精度(最坏0.82 ULP),GCC为2.8 ULP。MSVC的float erfcf(x)不仅速度慢2倍,精度也差一个数量级,甚至通过转换调用double erfc(x)都比原生erfcf(x)快2倍,十分反常。 - GCC优化需求:GCC编译的
dummy(x)函数(仅返回x)性能比其他编译器慢一个数量级,随机数据下表现更差,推测是编译器选项未设置得当导致求和溢出处理效率低。当前GCC编译命令为:gcc -O3 -march=native -mavx2 -finline-functions -Winline erfc_njaffa6.cpp -lm,寻求优化建议。 - 其他观察:GCC和Intel的float版本线性测试速度比double版本快约15%,符合预期,但GCC随机测试表现差。
内容的提问来源于stack exchange,提问作者Martin Brown
相关产品推荐
相关产品推荐

