NASM x86-64 DLL与ctypes交互:128位SIMD结果存取问题
1. Better SIMD Instructions for Storing 128-bit Results?
First, let's clarify why movdqa is crashing: movdqa requires the target memory to be 16-byte aligned. Even if you compiled with align=16, if the malloc-allocated buffer pointed to by rbp isn't properly aligned (e.g., pointer arithmetic shifted it off a 16-byte boundary, or your malloc implementation doesn't guarantee 16-byte alignment—though most x86-64 malloc implementations do), this will trigger an access violation.
Here are your practical options for storing the 128-bit result from xmm0:
- Stick with
movdquif alignment can't be guaranteed: It’s built for unaligned 128-bit memory access, so it won’t crash, and the performance difference is negligible for your 10k-loop workload. - Fix alignment to use
movdqa: Ensure the buffer is 16-byte aligned. On Windows, use_aligned_mallocinstead of standard malloc; on Linux/macOS, standard malloc already guarantees 16-byte alignment for x86-64. Double-check thatrbpisn’t offset by a non-multiple of 16 before the store operation. - Split into 64-bit stores with general-purpose instructions: If you want to avoid SIMD store instructions entirely, split the 128-bit result into two 64-bit values:
This works regardless of alignment but adds an extra instruction compared tomovq [rbp + rcx], xmm0 ; Store low 64 bits pextrq [rbp + rcx + 8], xmm0, 1 ; Store high 64 bits (index 1 for upper qword)movdqu.
2. Handling Returned 128-bit Integers in Python/ctypes
Since ctypes lacks a native 128-bit type, but Python supports arbitrary-precision integers, you can convert raw 16-byte chunks into usable Python integers with one of these approaches:
Option 1: Read as pairs of 64-bit integers and combine
Your buffer is a sequence of 16-byte blocks (each holding a 128-bit result). Cast the pointer to a c_uint64 pointer (matching your unsigned PMULUDQ usage), then combine pairs of 64-bit values into a single Python integer:
import ctypes # Assume CallName returns a pointer to your 128-bit result buffer result_ptr = CallName(...) n = 10000 # Number of 128-bit values # Cast to a pointer to unsigned 64-bit integers uint64_ptr = ctypes.cast(result_ptr, ctypes.POINTER(ctypes.c_uint64)) big_integers = [] for i in range(n): low = uint64_ptr[2 * i] # First 8 bytes: low 64 bits high = uint64_ptr[2 * i + 1] # Next 8 bytes: high 64 bits # Combine into a single 128-bit unsigned integer big_int = (high << 64) | low big_integers.append(big_int)
Option 2: Use a custom ctypes Structure for 128-bit values
Define a struct that represents a 128-bit integer as two 64-bit fields, then cast the result pointer to an array of this struct:
import ctypes class UInt128(ctypes.Structure): _fields_ = [ ("low", ctypes.c_uint64), ("high", ctypes.c_uint64) ] result_ptr = CallName(...) n = 10000 # Cast to a pointer to UInt128 array uint128_array = ctypes.cast(result_ptr, ctypes.POINTER(UInt128)) big_integers = [] for i in range(n): val = uint128_array[i] big_int = (val.high << 64) | val.low big_integers.append(big_int)
Option 3: Read raw bytes and convert with int.from_bytes
Read the entire buffer as a bytes object, then process each 16-byte chunk using Python's built-in int.from_bytes (x86-64 uses little-endian byte order):
import ctypes result_ptr = CallName(...) n = 10000 buffer_size = n * 16 # Read the entire buffer as bytes buffer = ctypes.string_at(result_ptr, buffer_size) big_integers = [] for i in range(n): # Extract each 16-byte chunk chunk = buffer[i*16 : (i+1)*16] # Convert to unsigned 128-bit integer (little-endian for x86-64) big_int = int.from_bytes(chunk, byteorder="little", signed=False) big_integers.append(big_int)
All three methods will give you Python arbitrary-precision integers that you can use directly in calculations.
内容的提问来源于stack exchange,提问作者RTC222

