固定尺寸二元多项式快速求值:MSVC与GCC编译优化对比
MSVC未展开std::array循环的原因与解决思路
我需要实现编译期固定尺寸的二元多项式快速求值,编写了如下示例代码:
#include <cmath> #include <array> #include <iostream> int main() { constexpr size_t NX = 5, NY = 4; using XA = std::array< double, NX >; using YA = std::array< XA, NY >; YA uu{}; for(size_t yi = 0; yi < NY; yi++) { for(size_t xi = 0; xi < NX; xi++) { uu[yi][xi] = xi + yi; std::cerr << "uu["<< yi << ","<< xi << "] = " << uu[yi][xi] << '\n'; } } double sum{0}, x = rand(), y = rand(); for(auto iy = uu.rbegin(); iy != uu.rend(); iy++) { auto ix = iy->rbegin(); double res = *ix++; for(; ix != iy->rend(); ix++) { res = std::fma(res, x, *ix); } sum = std::fma(sum, y, res); } std::cerr << "XXXX: " << sum << '\n'; return 0; }
使用GCC 8.1搭配-mfma -msse4.2 -O3 -DNDEBUG编译时,汇编代码已完全展开循环;但使用MSVC 2019搭配/W3 /GR /EHsc /Ox /MD /Ob2 /fp:fast /GL /arch:AVX2编译时,遍历std::array元素的循环并未展开。请问MSVC为何无法完成该循环展开?是否遗漏了某些优化选项?
原因分析
MSVC与GCC的循环展开逻辑存在差异,核心问题在于对std::array反向迭代器的常量推导能力:
- GCC能直接识别编译期固定大小的
std::array反向迭代器的遍历范围是常量值,因此自动触发完全展开; - MSVC 2019的优化器对
std::array反向迭代器的抽象层处理不够彻底,无法明确判断迭代器的遍历次数是固定值,因此不会主动展开循环。
你当前使用的/Ob2(内联扩展)虽然启用,但循环展开的触发依赖优化器对循环边界的确定性判断,反向迭代器的封装干扰了MSVC的这一判断。
优化选项与解决方法
强制循环展开
在循环前添加MSVC专用编译指令#pragma unroll,直接强制优化器展开循环,不受迭代器类型影响:// 外层循环 #pragma unroll for(auto iy = uu.rbegin(); iy != uu.rend(); iy++) { auto ix = iy->rbegin(); double res = *ix++; // 内层循环 #pragma unroll for(; ix != iy->rend(); ix++) { res = std::fma(res, x, *ix); } sum = std::fma(sum, y, res); }替换反向迭代器为下标遍历
改用基于编译期常量下标的遍历方式,让MSVC直接识别固定循环次数:double sum{0}, x = rand(), y = rand(); for(size_t yi = NY; yi > 0; ) { --yi; double res = uu[yi][NX-1]; for(size_t xi = NX-1; xi > 0; ) { --xi; res = std::fma(res, x, uu[yi][xi]); } sum = std::fma(sum, y, res); }验证链接器选项
你当前的编译选项已覆盖核心优化开关,没有遗漏关键项,但/GL(链接时优化)需要配合链接器选项/LTCG才能完全生效,确保编译时同时指定了该链接选项。
内容的提问来源于stack exchange,提问作者pem
相关产品推荐
相关产品推荐

