为何C++中使用+=的循环比等价=赋值循环运行速度更快?
问题现象
我在Visual Studio 2019 Release模式下编写了两段C++循环,均遍历长度100万的32位整型数组:
- 第一段为纯赋值逻辑,将每个元素赋值给它前一个位置的元素:
for ( int i = 1; i < Iters; i++ ) ints32[i - 1] = ints32[i];
- 第二段逻辑与前者完全一致,仅将赋值运算符替换为
+=:
for ( int i = 1; i < Iters; i++ ) ints32[i - 1] += ints32[i];
性能测试方法为:两个函数各运行20次,剔除5次最快、5次最慢的结果,对剩余10次结果取平均值,每次测试前都会将数组填充为随机整数。测试结果非常稳定:纯赋值循环耗时约460微秒,带加法的+=循环仅耗时320微秒,即便调换两个函数的测试顺序,该结果也保持不变。
按常理加法操作应该带来额外耗时,为什么带加法的循环反而更快?
两个函数的反汇编结果如下,除了加法循环多一条加法指令、eax初始指向ints[i-1]而非ints[i]之外,逻辑完全等价:
; 纯赋值循环 for ( int i = 1; i < Iters; i++ ) 00541670 B8 1C 87 54 00 mov eax,54871Ch ints32[i - 1] = ints32[i]; 00541675 8B 08 mov ecx,dword ptr [eax] 00541677 89 48 FC mov dword ptr [eax-4],ecx 0054167A 83 C0 04 add eax,4 0054167D 3D 18 90 91 00 cmp eax,offset floats32 (0919018h) 00541682 7C F1 jl IntAssignment32+5h (0541675h) } 00541684 C3 ret ; +=加法循环 for ( int i = 1; i < Iters; i++ ) 00541700 B8 18 87 54 00 mov eax,offset ints32 (0548718h) ints32[i - 1] += ints32[i]; 00541705 8B 08 mov ecx,dword ptr [eax] 00541707 03 48 04 add ecx,dword ptr [eax+4] 0054170A 89 08 mov dword ptr [eax],ecx 0054170C 83 C0 04 add eax,4 0054170F 3D 14 90 91 00 cmp eax,919014h 00541714 7C EF jl IntAddition32+5h (0541705h) } 00541716 C3 ret
注:整型数组被声明为
volatile,避免编译器做优化消除或者向量化等操作,生成的反汇编也符合测试预期。
另外修改程序中无关内容时,赋值版本的性能会出现明显波动,初步怀疑和代码对齐有关。
测试使用VS2019 Win32 Release默认编译选项:
/permissive- /ifcOutput "Release\" /GS /GL /analyze- /W3 /Gy /Zc:wchar_t /Zi /Gm- /O2 /sdl /Fd"Release\vc142.pdb" /Zc:inline /fp:precise /D "WIN32" /D "NDEBUG" /D "_CONSOLE" /D "_UNICODE" /D "UNICODE" /errorReport:prompt /WX- /Zc:forScope /Gd /Oy- /Oi /MD /FC /Fa"Release\" /EHsc /nologo /Fo"Release\" /Fp"Release\profile.pch" /diagnostics:column
编译器版本为Microsoft Visual Studio Community 2019 Version 16.11.15,完整测试代码见文末附录。
根本原因
核心原因就是你猜测的机器码地址对齐差异,和加法指令本身没有关系:
- x86处理器执行紧凑循环时,循环入口的内存地址是否对齐到16字节/32字节边界,会直接影响指令取指和译码的效率。如果关键的循环跳转目标落在地址不对齐的位置,哪怕只有1个字节的错位,也可能让每次循环迭代多消耗1个CPU时钟周期,100万次迭代累计下来就会形成非常明显的耗时差。
- 从你贴的反汇编可以直接看到两个循环的入口地址:
- 纯赋值循环的循环体入口(第一条
mov ecx,dword ptr [eax])地址是0x00541675 - 加法循环的循环体入口地址是
0x00541705
两个地址的对齐状态完全不同,这就是性能差异的直接来源。你修改无关代码时赋值版本性能波动,本质也是代码长度变化改变了循环入口的对齐位置,导致执行效率上下浮动。
- 纯赋值循环的循环体入口(第一条
- 可以做简单验证:给两个函数手动加对齐指令,或者在函数前插入几个
nop指令将循环入口调整到16字节对齐的位置,两个循环的耗时差就会消失,甚至纯赋值循环会因为少一条加法指令出现小幅的性能反超。 - 额外说明:因为你给数组加了
volatile限定,两个循环都完全没有做任何存储-加载转发优化,所有内存访问都严格按指令顺序执行,这种情况下前端取指的对齐开销占比会被大幅放大,才会出现“多一条指令反而更快”的反直觉现象。正常不加volatile的代码里,编译器会做更激进的指令调度和优化,这种对齐带来的性能差异会小很多。
附录:完整测试代码
#include <iostream> #include <chrono> #include <random> #include <algorithm> const int ValueRange = 100000000; std::default_random_engine generator; std::uniform_int_distribution< int > distribution( 1, ValueRange - 1 ); const int Iters = 1000000; // nanoseconds -> milliseconds volatile int ints32[Iters]; void InitArrayInt32() { for ( int i = 0; i < Iters; i++ ) ints32[i] = distribution( generator ); } const int SampleCount = 20; const int KeepSampleCount = SampleCount - 2 * (SampleCount / 4); float ProfileFunction( void(*setup)(), void(*func)() ) { uint64_t times[SampleCount]; for ( int i = 0; i < SampleCount; i++ ) { setup(); auto startTime = std::chrono::steady_clock::now(); func(); auto endTime = std::chrono::steady_clock::now(); times[i] = std::chrono::duration_cast<std::chrono::microseconds>( endTime - startTime ).count(); } std::sort( times, times + SampleCount ); uint64_t total = 0; for ( int i = SampleCount / 4; i < SampleCount - SampleCount / 4; i++ ) total += times[i]; return total * (1.0f / KeepSampleCount); } void IntAssignment32() { for ( int i = 1; i < Iters; i++ ) ints32[i - 1] = ints32[i]; } void IntAddition32() { for ( int i = 1; i < Iters; i++ ) ints32[i - 1] += ints32[i]; } int main() { float assignment = ProfileFunction( InitArrayInt32, IntAssignment32 ); float addition = ProfileFunction( InitArrayInt32, IntAddition32 ); printf( "assignment: %g\n", assignment ); printf( "addition: %g\n", addition ); return 0; }
内容的提问来源于stack exchange,提问作者John Haggerty
相关产品推荐
相关产品推荐

