ICX编译器中isnan()函数为何出现严重性能衰减?
Intel ICX 2023编译器性能异常衰减问题分析
问题复现场景
通过三次多项式Horner求值法逼近exp(x)的测试案例,对比两种系数存储布局的性能表现:
- stride=N的存储布局:性能表现正常
- stride=1的Horner b版本:出现严重性能衰减,相比MSC编译器性能下降4倍
性能衰减的核心诱因
- 函数内联失败:stride=1的Horner b版本无法被ICX编译器内联
- isnan()函数低效实现:该函数的实现会触发
sfence指令,导致流水线停滞 - 条件加法编译策略不合理:ICX对条件加法的编译逻辑进一步放大了性能损耗
汇编对比结论
通过对比不同编译器生成的汇编代码,已确认isnan()的低效实现、条件加法的不合理编译是性能差距的直接原因,但非内联代码中流水线停滞的根本原因尚未明确
可复现MRE代码示例
#include <cmath> #include <vector> constexpr int N = 1024; double coeffs_stride1[4] = {1.0, 1.0, 0.5, 1.0/6.0}; double coeffs_strideN[N*4]; // 初始化stride=N的系数数组 void init_coeffs() { for(int i=0; i<N; ++i) { coeffs_strideN[i*4 + 0] = 1.0; coeffs_strideN[i*4 + 1] = 1.0; coeffs_strideN[i*4 + 2] = 0.5; coeffs_strideN[i*4 + 3] = 1.0/6.0; } } // Horner求值法 - stride=N版本 double horner_a(const double* coeffs, double x) { return coeffs[0] + x*(coeffs[1] + x*(coeffs[2] + x*coeffs[3])); } // Horner求值法 - stride=1版本 double horner_b(const double* coeffs, double x) { double res = coeffs[3]; res = coeffs[2] + x*res; res = coeffs[1] + x*res; res = coeffs[0] + x*res; // 引入isnan()触发性能问题 if(std::isnan(res)) res = 0.0; return res; } // 性能测试入口 void test_perf() { init_coeffs(); double sum = 0.0; double x = 0.1; // 测试stride=N版本 for(int i=0; i<N; ++i) { sum += horner_a(&coeffs_strideN[i*4], x); } // 测试stride=1版本 for(int i=0; i<N; ++i) { sum += horner_b(coeffs_stride1, x); } }
内容的提问来源于stack exchange,提问作者Martin Brown
相关产品推荐
相关产品推荐

