在ASM中配置cmpxchg16b寄存器:C外部函数原子操作疑问
Alright, let's work through this together. You're trying to implement an atomic 128-bit compare-and-swap (CAS) using the lock cmpxchg16b instruction in an external assembly function for C, and you've got some questions about register handling—let's break this down step by step.
First, let's clear up a critical point: x86-64 registers are only 64 bits wide—you can't fit a full 128-bit value in a single register. Every 128-bit value needs two 64-bit registers (or a memory location) to hold its low and high halves. That's why your question about RDI having a "second half" makes sense—128-bit parameters are split across two registers in the standard x86-64 System V calling convention (used on Linux/macOS).
Step 1: Fix Parameter Mapping & C Function Declaration
Your initial parameter mapping (rdi: old status, rsi: cur status, rdx: mod status) needs adjustment because each 128-bit value occupies two register slots. For a standard atomic CAS operation (compare a memory value to an expected old value, swap with a new value if they match), here's the correct setup:
First, define the C function declaration that your assembly will implement:
// Returns 1 if the swap succeeded, 0 otherwise _Bool atomic_cas_128(__int128 *target_ptr, __int128 expected_old, __int128 new_mod);
Under the x86-64 System V convention, the parameters map to registers like this:
target_ptr(pointer to the 128-bit value we're modifying) →rdiexpected_old(128-bit): low 64 bits →rsi, high 64 bits →rdxnew_mod(128-bit): low 64 bits →r8, high 64 bits →r9
Step 2: Implement the lock cmpxchg16b Logic
The cmpxchg16b instruction operates on register pairs and memory:
- It compares the 128-bit value in
rax:rdx(low 64 inrax, high 64 inrdx) against the 128-bit value at the memory address inrdi - If they match, it writes the 128-bit value from
rbx:rcx(low 64 inrbx, high 64 inrcx) to that memory address and sets the ZF (zero flag) - If they don't match, it loads the current memory value into
rax:rdxand clears ZF
Here's the working assembly code (GNU assembler syntax):
.global atomic_cas_128 atomic_cas_128: // Load expected_old into rax:rdx (cmpxchg16b's expected value pair) movq %rsi, %rax // Low 64 bits of expected_old → rax // High 64 bits of expected_old are already in rdx (from calling convention) // Load new_mod into rbx:rcx (cmpxchg16b's replacement value pair) movq %r8, %rbx // Low 64 bits of new_mod → rbx movq %r9, %rcx // High 64 bits of new_mod → rcx // Execute atomic compare-and-swap on the target memory location lock cmpxchg16b (%rdi) // Compare rax:rdx with [rdi]; if match, write rbx:rcx to [rdi] // Set return value: 1 if swap succeeded (ZF set), 0 otherwise sete %al // Copy zero flag to 8-bit al register movzbl %al, %eax // Zero-extend al to 64-bit eax for proper return value ret
Step 3: Address Your Original Register Question
You asked about loading RDI and its "second half" into rax and rdx—that confusion comes from mixing up pointer vs. value parameters. In our setup, rdi holds a pointer to the 128-bit memory value we're modifying, not the value itself. The expected old value's halves are already passed in rsi (low) and rdx (high), so we only need to move the low half to rax to form the rax:rdx pair that cmpxchg16b expects.
Step 4: Call the Function from C
You can use the assembly function in C like this:
#include <stdint.h> // Declare the external assembly function _Bool atomic_cas_128(__int128 *target_ptr, __int128 expected_old, __int128 new_mod); int main() { __int128 my_val = ((__int128)0x123456789abcdef0 << 64) | 0x0fedcba987654321; __int128 expected = my_val; __int128 new_val = ((__int128)0xdeadbeefcafebabe << 64) | 0xbeefcafedeadbeef; _Bool success = atomic_cas_128(&my_val, expected, new_val); if (success) { // Swap worked—my_val is now new_val } else { // Swap failed—my_val was modified by another thread } return 0; }
Key Notes
- The
lockprefix is mandatory to makecmpxchg16batomic across multiple CPU cores. cmpxchg16bis supported on all x86-64 CPUs released since AMD Opteron and Intel Core 2.- If you're targeting Windows, you'll need to adjust the register mapping (Windows uses the fastcall convention, where the first parameter goes to
rcx, second tordx, etc.).
内容的提问来源于stack exchange,提问作者Shoisk

