You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为何GCC生成的浮点运算指令序列比Clang更快?

GCC vs Clang: Unexpected Floating-Point Performance Disparities

This is a fascinating example of how subtle compiler choices can lead to confusing performance observations—until you dig into the details of both the generated assembly and the benchmarking setup. Let's walk through what's happening here:

The Test Function

We're starting with this simple floating-point evaluation function:

float evaluate(float a, float b) { return (a - b + 1.0f) * (a - b) * (a - b - 1.0f); }

Generated Assembly (Optimized)

Both compilers were run with -std=c++1y -O2 flags. Here's the assembly each produced:

GCC 7.2 Output

evaluate(float, float):
 subss xmm0, xmm1
 movss xmm2, DWORD PTR .LC0[rip]
 movaps xmm1, xmm0
 addss xmm1, xmm2
 mulss xmm1, xmm0
 subss xmm0, xmm2
 mulss xmm0, xmm1
 ret
.LC0:
 .long 1065353216  # Encodes float 1.0f

Clang 5.0.0 Output

.LCPI0_0:
 .long 1065353216 # float 1
.LCPI0_1:
 .long 3212836864 # float -1
evaluate(float, float): # @evaluate(float, float)
 subss xmm0, xmm1
 movss xmm1, dword ptr [rip + .LCPI0_0] # xmm1 = mem[0],zero,zero,zero
 addss xmm1, xmm0
 mulss xmm1, xmm0
 addss xmm0, dword ptr [rip + .LCPI0_1]
 mulss xmm0, xmm1
 ret

Key Assembly Differences

Breaking down the two outputs, we can see several deliberate compiler choices:

  • GCC uses movaps instead of movss when copying xmm0 to xmm1. As Peter Cordes noted, this avoids stalls caused by partial updates to XMM registers—even though movss would technically work for this scalar operation.
  • GCC leverages 3 XMM registers, while Clang sticks to just 2.
  • Clang stores both 1.0f and -1.0f as constants, using addss for both the +1 and -1 operations. GCC only stores 1.0f, using subss to handle the -1 step instead.

The Performance Paradox

When mapping out instruction dependency chains, Clang's code looked like it should be faster—it had a shorter dependency path. But initial benchmarks told a different story: GCC's version was roughly 30% faster on both Intel Broadwell and AMD Bulldozer CPUs.

Resolving the Discrepancy

The root cause turned out to be a flawed benchmark setup: the initial test used low-quality inline assembly that introduced unintended overhead. Once we switched to compiling each version to separate object files and linking them together, the performance difference completely disappeared.

Peter Cordes hypothesized that the initial gap might have been due to -O1 omitting .p2align directives, and confirmed that both implementations matched the assembly listings above once the benchmark was fixed.

内容的提问来源于stack exchange,提问作者Bernardo Sulzbach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:25:09