You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

x86架构下比特打包时Shift与Add指令的性能差异探究

移位与Add+位运算实现比特打包的性能差异疑惑

我测试了一段C代码,实现两种int8_t数组的比特打包逻辑:一种用Shift指令移位组合,另一种用Add与位运算替代移位。多次编译测试显示,无Shift版本性能更优。查看Godbolt生成的汇编,Shift版本对应sal指令,无Shift版本对应add+and指令。但根据Agner Fog的指令性能表,sal和两条add指令理论性能相近,想不通为什么实测有明显差异。


测试代码

#include <stdio.h>
#include <inttypes.h>
#include <time.h>

#define NANOSEC_PER_SEC 1000000000LL

static inline void PackInt8Shifts(const int8_t* values, volatile uint8_t* out) {
  *out = (values[0] | values[1] << 1 | values[2] << 2 | values[3] << 3 | values[4] << 4 |
          values[5] << 5 | values[6] << 6 | values[7] << 7);
}

static inline void PackInt8NoShifts(const int8_t* values, volatile uint8_t* out) {
  *out = (values[0] | ((values[1] + 0x1) & 0x2) | ((values[2] + 0x3) & 0x4) |
          ((values[3] + 0x7) & 0x8) | ((values[4] + 0xf) & 0x10) |
          ((values[5] + 0x1f) & 0x20) | ((values[6] + 0x3f) & 0x40) |
          ((values[7] + 0x7f) & 0x80));
}


int main(void) {
  const size_t niters = 100000000;
  const int8_t values[8] = {1, 1, 0, 1, 0, 0, 1, 1};
  volatile uint8_t out;
  struct timespec start, end;
  
  clock_gettime(CLOCK_REALTIME, &start);
  for (size_t i = 0; i < niters; i++) {
    PackInt8Shifts(values, &out);
  }
  clock_gettime(CLOCK_REALTIME, &end);
  printf("ns duration of PackInt8Shifts was: %lld\n",
         (end.tv_sec * NANOSEC_PER_SEC + end.tv_nsec)
         - (start.tv_sec * NANOSEC_PER_SEC + start.tv_nsec));

  clock_gettime(CLOCK_REALTIME, &start);
  for (size_t i = 0; i < niters; i++) {
    PackInt8NoShifts(values, &out);
  }
  clock_gettime(CLOCK_REALTIME, &end);
  printf("ns duration of PackInt8NoShifts was: %lld\n",
         (end.tv_sec * NANOSEC_PER_SEC + end.tv_nsec)
         - (start.tv_sec * NANOSEC_PER_SEC + start.tv_nsec));

  return 0;
}

多次测试结果

$ gcc -O3 packint.c 
$ ./a.out
ns duration of PackInt8Shifts was: 17310031
ns duration of PackInt8NoShifts was: 11727328
$ ./a.out
ns duration of PackInt8Shifts was: 16955954
ns duration of PackInt8NoShifts was: 11287199
$ ./a.out
ns duration of PackInt8Shifts was: 17187124
ns duration of PackInt8NoShifts was: 11239957

对应汇编代码

Shift版本

mov     rax, QWORD PTR [rbp-8]
        add     rax, 2
        movzx   eax, BYTE PTR [rax]
        sal     eax, 2
        or      edx, eax

无Shift版本

mov     rax, QWORD PTR [rbp-8]
        add     rax, 2
        movzx   eax, BYTE PTR [rax]
        add     eax, 3
        and     eax, 4
        or      edx, eax

Agner Fog指令性能表

InstructionOperandsOpsLatencyReciprocal Throughput
ADD, SUBr,i110.25
SHL, SHR, SARr,i/CL110.5

性能差异的核心原因

  1. 执行单元资源竞争:ADD指令可以在CPU多个通用整数执行单元(如Intel CPU的端口0、1、5)并行执行,而SAL移位指令通常只能占用特定的移位执行单元(如端口1)。你的打包操作包含多次同类指令,无Shift版本的ADD+AND能分散到更多单元并行处理,Shift版本的SAL则会集中占用有限的移位单元,形成执行瓶颈。
  2. 指令级并行度限制:单条指令延迟相同,但无Shift版本的指令依赖链更松散,CPU乱序执行引擎更容易调度多个操作并行。比如ADD的结果可立即被AND使用,而移位指令的结果直接参与后续OR,依赖关系更紧密,限制了并行空间。
  3. 微码优化差异:无Shift版本的+常量+&常量操作,编译器可能将其融合为更高效的微码执行;而移位指令的微码实现相对复杂,即使是固定移位量,执行效率也不如分散的加法与位运算组合。

注:测试中的volatile变量强制每次循环写内存,但两种版本的内存操作逻辑一致,性能差异完全来自计算部分的指令执行效率。

内容的提问来源于stack exchange,提问作者Will Ayd

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 12:22:16