std::atomic非无锁实现的锁存储位置及多实例影响问询
Great question! Let's unpack this clearly, using your example and GCC's implementation (since you mentioned a Linux/gcc environment) as a reference.
First, your observation is spot-on: since std::atomic<foo> takes up the exact same 16 bytes as foo itself, the lock can't be stored inside the atomic variable. So where is it?
Where the lock lives
When std::atomic<T> can't be lock-free (i.e., is_lock_free() returns 0), the standard library (like GCC's libstdc++) uses a global/static lock pool instead of attaching a lock to each atomic instance. Here's how it works:
- The library maintains a fixed set of global locks. When you perform an atomic operation on a non-lock-free
std::atomic<T>object, it hashes the object's memory address to pick one lock from this pool. - This design avoids adding extra memory overhead to each atomic instance (which explains why
sizeof(var)equalssizeof(foo)in your example) and prevents wasting memory on thousands of unused locks if you have many atomic objects.
Impact on multiple instances
This shared lock pool has two key implications for multiple std::atomic<foo> instances:
- Reduced concurrency: If multiple atomic objects hash to the same lock, operations on any of them will block operations on the others. For example, if you create 100
std::atomic<foo>objects, several might share a single lock—so concurrent access to these objects will cause unnecessary blocking, killing the parallelism you'd expect from atomic types. - Lost lock-free benefits: Lock-free atomics shine because they avoid lock contention and overhead. Once you're using a lock-based implementation, you lose those advantages; performance will be roughly equivalent to protecting a regular
fooobject with a standard mutex.
A quick side note
You might be wondering: x86_64 supports the cmpxchg16b instruction, which can atomically manipulate 16-byte values—so why does is_lock_free() return 0 here? Lock-free support for std::atomic<T> depends on both hardware capabilities and compiler implementation. In some cases, even if the hardware can handle it, the compiler may not implement lock-free operations for your specific type (e.g., due to alignment requirements or library design choices), forcing a fallback to the lock-based approach.
内容的提问来源于stack exchange,提问作者curiousguy12

