You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在ASM中配置cmpxchg16b寄存器:C外部函数原子操作疑问

Alright, let's work through this together. You're trying to implement an atomic 128-bit compare-and-swap (CAS) using the lock cmpxchg16b instruction in an external assembly function for C, and you've got some questions about register handling—let's break this down step by step.

First, let's clear up a critical point: x86-64 registers are only 64 bits wide—you can't fit a full 128-bit value in a single register. Every 128-bit value needs two 64-bit registers (or a memory location) to hold its low and high halves. That's why your question about RDI having a "second half" makes sense—128-bit parameters are split across two registers in the standard x86-64 System V calling convention (used on Linux/macOS).

Step 1: Fix Parameter Mapping & C Function Declaration

Your initial parameter mapping (rdi: old status, rsi: cur status, rdx: mod status) needs adjustment because each 128-bit value occupies two register slots. For a standard atomic CAS operation (compare a memory value to an expected old value, swap with a new value if they match), here's the correct setup:

First, define the C function declaration that your assembly will implement:

// Returns 1 if the swap succeeded, 0 otherwise
_Bool atomic_cas_128(__int128 *target_ptr, __int128 expected_old, __int128 new_mod);

Under the x86-64 System V convention, the parameters map to registers like this:

  • target_ptr (pointer to the 128-bit value we're modifying) → rdi
  • expected_old (128-bit): low 64 bits → rsi, high 64 bits → rdx
  • new_mod (128-bit): low 64 bits → r8, high 64 bits → r9

Step 2: Implement the lock cmpxchg16b Logic

The cmpxchg16b instruction operates on register pairs and memory:

  • It compares the 128-bit value in rax:rdx (low 64 in rax, high 64 in rdx) against the 128-bit value at the memory address in rdi
  • If they match, it writes the 128-bit value from rbx:rcx (low 64 in rbx, high 64 in rcx) to that memory address and sets the ZF (zero flag)
  • If they don't match, it loads the current memory value into rax:rdx and clears ZF

Here's the working assembly code (GNU assembler syntax):

.global atomic_cas_128
atomic_cas_128:
    // Load expected_old into rax:rdx (cmpxchg16b's expected value pair)
    movq    %rsi, %rax        // Low 64 bits of expected_old → rax
    // High 64 bits of expected_old are already in rdx (from calling convention)

    // Load new_mod into rbx:rcx (cmpxchg16b's replacement value pair)
    movq    %r8, %rbx         // Low 64 bits of new_mod → rbx
    movq    %r9, %rcx         // High 64 bits of new_mod → rcx

    // Execute atomic compare-and-swap on the target memory location
    lock cmpxchg16b (%rdi)    // Compare rax:rdx with [rdi]; if match, write rbx:rcx to [rdi]

    // Set return value: 1 if swap succeeded (ZF set), 0 otherwise
    sete    %al               // Copy zero flag to 8-bit al register
    movzbl  %al, %eax         // Zero-extend al to 64-bit eax for proper return value
    ret

Step 3: Address Your Original Register Question

You asked about loading RDI and its "second half" into rax and rdx—that confusion comes from mixing up pointer vs. value parameters. In our setup, rdi holds a pointer to the 128-bit memory value we're modifying, not the value itself. The expected old value's halves are already passed in rsi (low) and rdx (high), so we only need to move the low half to rax to form the rax:rdx pair that cmpxchg16b expects.

Step 4: Call the Function from C

You can use the assembly function in C like this:

#include <stdint.h>

// Declare the external assembly function
_Bool atomic_cas_128(__int128 *target_ptr, __int128 expected_old, __int128 new_mod);

int main() {
    __int128 my_val = ((__int128)0x123456789abcdef0 << 64) | 0x0fedcba987654321;
    __int128 expected = my_val;
    __int128 new_val = ((__int128)0xdeadbeefcafebabe << 64) | 0xbeefcafedeadbeef;

    _Bool success = atomic_cas_128(&my_val, expected, new_val);
    if (success) {
        // Swap worked—my_val is now new_val
    } else {
        // Swap failed—my_val was modified by another thread
    }
    return 0;
}

Key Notes

  • The lock prefix is mandatory to make cmpxchg16b atomic across multiple CPU cores.
  • cmpxchg16b is supported on all x86-64 CPUs released since AMD Opteron and Intel Core 2.
  • If you're targeting Windows, you'll need to adjust the register mapping (Windows uses the fastcall convention, where the first parameter goes to rcx, second to rdx, etc.).

内容的提问来源于stack exchange,提问作者Shoisk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 03:44:39