为何GCC生成的浮点运算指令序列比Clang更快?
This is a fascinating example of how subtle compiler choices can lead to confusing performance observations—until you dig into the details of both the generated assembly and the benchmarking setup. Let's walk through what's happening here:
The Test Function
We're starting with this simple floating-point evaluation function:
float evaluate(float a, float b) { return (a - b + 1.0f) * (a - b) * (a - b - 1.0f); }
Generated Assembly (Optimized)
Both compilers were run with -std=c++1y -O2 flags. Here's the assembly each produced:
GCC 7.2 Output
evaluate(float, float): subss xmm0, xmm1 movss xmm2, DWORD PTR .LC0[rip] movaps xmm1, xmm0 addss xmm1, xmm2 mulss xmm1, xmm0 subss xmm0, xmm2 mulss xmm0, xmm1 ret .LC0: .long 1065353216 # Encodes float 1.0f
Clang 5.0.0 Output
.LCPI0_0: .long 1065353216 # float 1 .LCPI0_1: .long 3212836864 # float -1 evaluate(float, float): # @evaluate(float, float) subss xmm0, xmm1 movss xmm1, dword ptr [rip + .LCPI0_0] # xmm1 = mem[0],zero,zero,zero addss xmm1, xmm0 mulss xmm1, xmm0 addss xmm0, dword ptr [rip + .LCPI0_1] mulss xmm0, xmm1 ret
Key Assembly Differences
Breaking down the two outputs, we can see several deliberate compiler choices:
- GCC uses
movapsinstead ofmovsswhen copyingxmm0toxmm1. As Peter Cordes noted, this avoids stalls caused by partial updates to XMM registers—even thoughmovsswould technically work for this scalar operation. - GCC leverages 3 XMM registers, while Clang sticks to just 2.
- Clang stores both
1.0fand-1.0fas constants, usingaddssfor both the+1and-1operations. GCC only stores1.0f, usingsubssto handle the-1step instead.
The Performance Paradox
When mapping out instruction dependency chains, Clang's code looked like it should be faster—it had a shorter dependency path. But initial benchmarks told a different story: GCC's version was roughly 30% faster on both Intel Broadwell and AMD Bulldozer CPUs.
Resolving the Discrepancy
The root cause turned out to be a flawed benchmark setup: the initial test used low-quality inline assembly that introduced unintended overhead. Once we switched to compiling each version to separate object files and linking them together, the performance difference completely disappeared.
Peter Cordes hypothesized that the initial gap might have been due to -O1 omitting .p2align directives, and confirmed that both implementations matched the assembly listings above once the benchmark was fixed.
内容的提问来源于stack exchange,提问作者Bernardo Sulzbach

