x86架构下比特打包时Shift与Add指令的性能差异探究
移位与Add+位运算实现比特打包的性能差异疑惑
我测试了一段C代码,实现两种int8_t数组的比特打包逻辑:一种用Shift指令移位组合,另一种用Add与位运算替代移位。多次编译测试显示,无Shift版本性能更优。查看Godbolt生成的汇编,Shift版本对应sal指令,无Shift版本对应add+and指令。但根据Agner Fog的指令性能表,sal和两条add指令理论性能相近,想不通为什么实测有明显差异。
测试代码
#include <stdio.h> #include <inttypes.h> #include <time.h> #define NANOSEC_PER_SEC 1000000000LL static inline void PackInt8Shifts(const int8_t* values, volatile uint8_t* out) { *out = (values[0] | values[1] << 1 | values[2] << 2 | values[3] << 3 | values[4] << 4 | values[5] << 5 | values[6] << 6 | values[7] << 7); } static inline void PackInt8NoShifts(const int8_t* values, volatile uint8_t* out) { *out = (values[0] | ((values[1] + 0x1) & 0x2) | ((values[2] + 0x3) & 0x4) | ((values[3] + 0x7) & 0x8) | ((values[4] + 0xf) & 0x10) | ((values[5] + 0x1f) & 0x20) | ((values[6] + 0x3f) & 0x40) | ((values[7] + 0x7f) & 0x80)); } int main(void) { const size_t niters = 100000000; const int8_t values[8] = {1, 1, 0, 1, 0, 0, 1, 1}; volatile uint8_t out; struct timespec start, end; clock_gettime(CLOCK_REALTIME, &start); for (size_t i = 0; i < niters; i++) { PackInt8Shifts(values, &out); } clock_gettime(CLOCK_REALTIME, &end); printf("ns duration of PackInt8Shifts was: %lld\n", (end.tv_sec * NANOSEC_PER_SEC + end.tv_nsec) - (start.tv_sec * NANOSEC_PER_SEC + start.tv_nsec)); clock_gettime(CLOCK_REALTIME, &start); for (size_t i = 0; i < niters; i++) { PackInt8NoShifts(values, &out); } clock_gettime(CLOCK_REALTIME, &end); printf("ns duration of PackInt8NoShifts was: %lld\n", (end.tv_sec * NANOSEC_PER_SEC + end.tv_nsec) - (start.tv_sec * NANOSEC_PER_SEC + start.tv_nsec)); return 0; }
多次测试结果
$ gcc -O3 packint.c $ ./a.out ns duration of PackInt8Shifts was: 17310031 ns duration of PackInt8NoShifts was: 11727328 $ ./a.out ns duration of PackInt8Shifts was: 16955954 ns duration of PackInt8NoShifts was: 11287199 $ ./a.out ns duration of PackInt8Shifts was: 17187124 ns duration of PackInt8NoShifts was: 11239957
对应汇编代码
Shift版本
mov rax, QWORD PTR [rbp-8] add rax, 2 movzx eax, BYTE PTR [rax] sal eax, 2 or edx, eax
无Shift版本
mov rax, QWORD PTR [rbp-8] add rax, 2 movzx eax, BYTE PTR [rax] add eax, 3 and eax, 4 or edx, eax
Agner Fog指令性能表
| Instruction | Operands | Ops | Latency | Reciprocal Throughput |
|---|---|---|---|---|
| ADD, SUB | r,i | 1 | 1 | 0.25 |
| SHL, SHR, SAR | r,i/CL | 1 | 1 | 0.5 |
性能差异的核心原因
- 执行单元资源竞争:
ADD指令可以在CPU多个通用整数执行单元(如Intel CPU的端口0、1、5)并行执行,而SAL移位指令通常只能占用特定的移位执行单元(如端口1)。你的打包操作包含多次同类指令,无Shift版本的ADD+AND能分散到更多单元并行处理,Shift版本的SAL则会集中占用有限的移位单元,形成执行瓶颈。 - 指令级并行度限制:单条指令延迟相同,但无Shift版本的指令依赖链更松散,CPU乱序执行引擎更容易调度多个操作并行。比如
ADD的结果可立即被AND使用,而移位指令的结果直接参与后续OR,依赖关系更紧密,限制了并行空间。 - 微码优化差异:无Shift版本的
+常量+&常量操作,编译器可能将其融合为更高效的微码执行;而移位指令的微码实现相对复杂,即使是固定移位量,执行效率也不如分散的加法与位运算组合。
注:测试中的volatile变量强制每次循环写内存,但两种版本的内存操作逻辑一致,性能差异完全来自计算部分的指令执行效率。
内容的提问来源于stack exchange,提问作者Will Ayd
相关产品推荐
相关产品推荐

