You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

MSVC中float erfcf(x)比double erfc(x)慢2倍?求复现及GCC优化建议

关于MSVC 17.1中erfcf(x)性能异常的验证与GCC优化建议

我需要验证在MSVC 17.1标准库中观察到的一个性能异常:float erfcf(x)的速度比double erfc(x)慢2倍,确认这是普遍现象而非基准测试偏差;同时寻求GCC编译代码时的优化命令行参数建议。

我采用@Njuffa的测试框架对erfcf(x)函数开展实验,为跟踪进度添加了进度打印,发现线性遍历x的所有位模式时,打印会在特定值范围停滞,以此指导优化。测试中Intel 2024.1 ICX库的实现兼具速度与精度,但MSVC存在明显性能异常,且GCC 13.1编译的dummy循环测试速度极慢,怀疑是编译器选项设置或虚拟机环境导致,希望得到原生Linux系统的基准测试数据对比。

最小可复现代码

#include <stdio.h>
#include <stdint.h>
#include <string.h>
#include <time.h>
#include <math.h>

//#define RANDOM  (1)

// helper routines

float uint32_as_float(uint32_t a)
{
   float r;
   memcpy(&r, &a, sizeof r);
   return r;
}

uint32_t float_as_uint32(float a)
{
   uint32_t r;
   memcpy(&r, &a, sizeof r);
   return r;
}

bool my_isnan(float x)
{
  //    used here so that -ffast-math can't optimise it away (as it does with system isnan()
  //    other options can be used here
  //    return __isnan((double)x);
  //    return __isnanf(x);
  uint32_t ix = float_as_uint32(x);
  //    return !((~ix) & 0x7f800000);  // Intel's choice is slightly faster
  return (ix & 0x7f800000) == 0x7f800000;
}

float erfcdf(float x)
{
  return (float)erfc((double)x);  // shim for calling double precision erfc
}

float dummy(float x)  
{
  return x;  // to determine loop overheads
}

void timefun(const char* name, float (*test_fun)(float), bool verbose)
{
  uint32_t argi, largi = 0;
  float arg, res, sum;
  time_t start0, start, end;
  printf("\nTiming %s\n", name);
  argi = 0;
  sum = 0.0;
  start0 = start = clock();
  do {
      arg = uint32_as_float(argi);
      res = (*test_fun)(arg);
      if (!my_isnan(res)) sum += res;
#ifdef RANDOM
      argi = (argi * 1664525 + 1013904223); // ranqd1 
#else
      argi++;
#endif
      if (verbose && ((argi & 0xff800000) != largi))
      {
          end = clock();
          largi = argi & 0xff800000;
          printf("Exp %x : %6.3f\n", argi >> 20, (float)(end - start) / CLOCKS_PER_SEC);
          start = clock();
      }
      if ((argi & 0x3ffffff) == 0) printf(".");
  } while (argi);
  end = clock();
  printf("\ntime taken %6.2f  sum = %g\n", (float)(end - start0) / CLOCKS_PER_SEC, sum);
}

int main(void)
{
  timefun("dummy", dummy, false);
  timefun("erfcf", erfcf, false);
  timefun("erfcdf", erfcdf, false);
  return 0;
}

该代码可在GCC、Intel、MSVC编译器上直接编译,也欢迎其他编译器的基准测试结果。

基准测试数据

(线性测试遍历x的所有位模式,随机测试用ranqd1打乱分支预测,单位:秒)

编译器Dummy线性测试线性erfcf线性erfc随机Dummy随机erfcf随机erfc
gcc 13.13131.651.138.7192169
Intel 2024.12.037.645.03.946.954.3
MS 17.12.096.247.64.0275125

分析与需求

  • MSVC异常分析:推测性能慢的原因是对非正规数(denorms)的多项式处理过于简单,导致速度骤降;精度测试显示,MSVC的float erfcf(x)仅达5.7 ULP,而所有编译器的double erfc(x)转为float后精度为0.5 ULP,Intel的float erfcf(x)更是达到亚ULP精度(最坏0.82 ULP),GCC为2.8 ULP。MSVC的float erfcf(x)不仅速度慢2倍,精度也差一个数量级,甚至通过转换调用double erfc(x)都比原生erfcf(x)快2倍,十分反常。
  • GCC优化需求:GCC编译的dummy(x)函数(仅返回x)性能比其他编译器慢一个数量级,随机数据下表现更差,推测是编译器选项未设置得当导致求和溢出处理效率低。当前GCC编译命令为:gcc -O3 -march=native -mavx2 -finline-functions -Winline erfc_njaffa6.cpp -lm,寻求优化建议。
  • 其他观察:GCC和Intel的float版本线性测试速度比double版本快约15%,符合预期,但GCC随机测试表现差。

内容的提问来源于stack exchange,提问作者Martin Brown

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 08:43:12