为何C#中(result >> 8) & hfm比((operand1^result^operand2)>>8)&hfm更慢?
位运算代码Debug模式下的性能反直觉现象
测试代码
void Main() { const byte hfm = 0b0001_0000; UInt16 operand1 = 0x2; UInt16 operand2 = 0xFFFE; UInt32 result = (UInt32)(operand1 + operand2); UInt32 hf; var timer = new Stopwatch(); timer.Start(); for (var i = 0; i < 500_000_000; i++) { hf = (((operand1 ^ result ^ operand2) >> 8) & hfm); } timer.Stop(); timer.Elapsed.Dump("(((operand1 ^ result ^ operand2) >> 8) & hfm)"); timer.Restart(); for (var i = 0; i < 500_000_000; i++) { hf = (result >> 8) & hfm; } timer.Stop(); timer.Elapsed.Dump("(result >> 8) & hfm"); }
Debug模式测试结果
(((operand1 ^ result ^ operand2) >> 8) & hfm) 00:00:01.1278598 (result >> 8) & hfm 00:00:01.1612988
代码更简洁的第二段循环,耗时反而比第一段多约0.03秒。
逐步简化后的性能变化
- 移除两处的
& hfm后,第二段代码耗时基本不变,第一段仅比第二段快约0.002秒 - 再移除
>> 8后,第一段才比第二段快约0.01秒
对应的IL指令
第一段代码的IL
IL_0022 ldloc.0 // operand1 IL_0023 ldloc.2 // result IL_0024 xor IL_0025 ldloc.1 // operand2 IL_0026 xor IL_0027 ldc.i4.8 IL_0028 shr.un IL_0029 ldc.i4.s 10 // 16 IL_002B and IL_002C stloc.3 // hf
第二段代码的IL
IL_0073 ldloc.2 // result IL_0074 ldc.i4.8 IL_0075 shr.un IL_0076 ldc.i4.s 10 // 16 IL_0078 and IL_0079 stloc.3 // hf
从IL指令来看,第二段的操作步骤更少,理论上应该更快,但实际测试结果相反。
最终结论
进一步测试发现,LINQPad默认以Debug模式编译代码,切换为Release模式后两段代码的性能基本相近。Debug模式下的性能差异源于该模式会禁用大部分优化逻辑,同时插入调试相关的辅助代码,导致实际执行效率和IL层面的理论分析出现偏差。
内容的提问来源于stack exchange,提问作者Lee
相关产品推荐
相关产品推荐

