You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为什么g++开启O2优化后无符号整数回绕特性无法正常工作?

问题复现
  • 参考Intel DPDK的libring组件,利用无符号整数回绕特性实现C++队列,代码如下:
#include <cstdio>
#include <cassert>
#include <atomic>
#include <thread>

size_t global_r = 0, global_w = 0, mask_ = 3;

void emplace() {
  unsigned long local_w, local_r, free_entries = 0;
  local_w = global_w;
  while (free_entries == 0) {
    local_r = global_r;
    free_entries = (mask_ + local_r - local_w);
  }
  fprintf(stderr, "%lu\n", free_entries);
  auto w_next = local_w + 1;
  std::atomic_thread_fence(std::memory_order_release);
  global_w = w_next;
}

void pop() {
  unsigned long local_r = global_r;
  unsigned long r_next = local_r + 1;
  // make sure nobody can write to it before destruction
  std::atomic_thread_fence(std::memory_order_release);
  global_r = r_next;
}

int main() {
  std::jthread([]() -> void {
    int i = 10;
    while (i-- >= 0) emplace();
  });
  std::jthread([]() -> void {
    int i = 10;
    while (i-- >= 0) pop();
  });
  return 0;
}
  • 不同优化等级下运行结果差异极大:
    1. O2优化结果:
    3
    2
    1
    0
    18446744073709551615
    18446744073709551614
    18446744073709551613
    18446744073709551612
    18446744073709551611
    18446744073709551610
    18446744073709551609
    
    1. O0优化结果:
    3
    2
    1
    .....long time suspending
    
  • 核心疑问:无符号整数回绕是C++标准明确规定的合法行为,当前问题是对该特性理解有偏差,还是代码存在其他问题?

问题原因
  1. 核心问题和无符号整数回绕特性无关,回绕行为本身是符合标准的,问题出在多线程共享变量的访问规则违反了C++内存模型要求。
  2. 全局变量global_r、global_w、mask_都不是原子类型,多线程并发读写非原子共享变量属于未定义行为,编译器有权做任何优化:
    • O0优化等级下,编译器不会做激进的寄存器缓存,每次访问变量都会直接读写内存,但因为没有同步操作,写线程读取到的global_r始终没有更新,会一直卡在free_entries == 0的循环中,表现为程序挂起。
    • O2优化等级下,编译器判定非原子变量global_r不会被其他线程修改,直接将local_r = global_r的读取操作提到循环外部,循环内不会再重新读取内存中的global_r最新值,local_w持续递增时,mask_ + local_r - local_w计算结果会不断减小,出现负数后触发无符号回绕,就会输出你看到的超大数值。
  3. 单独加std::atomic_thread_fence不会生效,内存栅栏的同步效果需要和原子变量的读写搭配使用,对非原子变量的读写没有同步约束。
修复方案

将所有跨线程访问的共享变量改为原子类型,搭配正确的内存序读写即可:

#include <cstdio>
#include <atomic>
#include <thread>

std::atomic<size_t> global_r = 0, global_w = 0;
constexpr size_t mask_ = 3;

void emplace() {
  size_t local_w, local_r, free_entries = 0;
  local_w = global_w.load(std::memory_order_relaxed);
  while (free_entries == 0) {
    local_r = global_r.load(std::memory_order_acquire);
    free_entries = (mask_ + local_r - local_w);
  }
  fprintf(stderr, "%lu\n", free_entries);
  auto w_next = local_w + 1;
  global_w.store(w_next, std::memory_order_release);
}

void pop() {
  size_t local_r = global_r.load(std::memory_order_relaxed);
  size_t r_next = local_r + 1;
  global_r.store(r_next, std::memory_order_release);
}

int main() {
  std::jthread t1([]() -> void {
    int i = 10;
    while (i-- >= 0) emplace();
  });
  std::jthread t2([]() -> void {
    int i = 10;
    while (i-- >= 0) pop();
  });
  return 0;
}

内容的提问来源于stack exchange,提问作者Pan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 08:06:02