关于Thread-local与C++11内存模型中load-acquire/store-release同步的咨询
Hey there, let’s dive into these synchronization concepts—they’re key to writing correct concurrent C++ code, so it’s great you’re digging into the details!
First: C++11’s Load-Acquire/Store-Release Basics
Let’s start with the core mechanism you mentioned, since it’s the foundation for most safe concurrent operations (including how mutexes work under the hood).
At their core, load(std::memory_order_acquire) and store(std::memory_order_release) enforce two critical things across threads: visibility and ordering:
- Load-Acquire: When a thread executes this load, it guarantees two things:
- No subsequent memory operations (reads or writes) in the thread will be reordered before this load.
- The load will see the results of all prior
store(std::memory_order_release)operations on the same atomic variable from other threads.
- Store-Release: When a thread executes this store, it guarantees:
- No prior memory operations in the thread will be reordered after this store.
- The results of all those prior operations will be visible to any thread that later does a
load(std::memory_order_acquire)on the same atomic variable.
This is exactly how mutexes work: calling lock() on a mutex is an acquire operation, and unlock() is a release operation. Any code between lock() and unlock() (the "critical section") is guaranteed to be synchronized across threads—all observing threads will agree on the order of operations inside that section.
Read-Modify-Write (RMW) Operations
These acquire/release semantics often tie into read-modify-write instructions (like std::atomic_fetch_add, std::atomic_exchange), which can use std::memory_order_acq_rel to get both acquire and release behavior in a single operation. This is how you build lock-free structures (like spinlocks):
std::atomic<bool> spin_lock(false); // Thread A: Acquire the lock while (spin_lock.exchange(true, std::memory_order_acquire)) { // Spin until lock is free } // Critical section: All operations here won't be reordered outside the lock spin_lock.store(false, std::memory_order_release); // Thread B: Same logic while (spin_lock.exchange(true, std::memory_order_acquire)) {} // Critical section operations are synchronized with Thread A's spin_lock.store(false, std::memory_order_release);
In this example, the exchange (a RMW operation) uses acquire to ensure critical section code doesn’t sneak before the lock acquisition, and the final store(release) ensures all critical section operations are visible to the next thread that acquires the lock.
Thread-local Acquire/Release Synchronization
Now, thread-local storage (TLS) is a bit special because each thread gets its own private copy of the variable. Most of the time, you don’t need synchronization for TLS variables—since no other thread can access your thread’s copy directly.
But if you do share a pointer to a thread-local variable across threads (e.g., passing a pointer to your TLS atomic to another thread), then acquire/release semantics still apply exactly as they do for regular atomics.
For example:
// Each thread has its own copy of this atomic thread_local std::atomic<int> thread_state(0); // Thread 1: Update its TLS state and signal Thread 2 thread_state.store(42, std::memory_order_release); // Pass &thread_state to Thread 2 somehow (e.g., via a queue) // Thread 2: Access Thread 1's TLS atomic int state = (*thread1_state_ptr).load(std::memory_order_acquire); // Thread 2 is guaranteed to see state = 42, and all operations Thread 1 did before the store
The key takeaway here is: TLS doesn’t change the rules of acquire/release—it just changes who has access to the variable by default. Only when you break that "thread-private" boundary do you need to apply synchronization.
内容的提问来源于stack exchange,提问作者Siler

