如何用NEON高效转置以四个int32x4t表示的4×4 int32矩阵?
高效转置4x4 int32 NEON矩阵(禁用ld4q_s32/st4q_s32)
针对由四个int32x4t向量表示的4x4 int32矩阵,可通过NEON的向量拆分、交错重组指令实现高效转置,完全避开ld4q_s32和st4q_s32。以下是具体实现:
实现思路
输入的四个int32x4t对应矩阵的四行:
Row0: a0 a1 a2 a3 Row1: b0 b1 b2 b3 Row2: c0 c1 c2 c3 Row3: d0 d1 d2 d3
转置目标是将列转为行,得到四个向量:
Col0: a0 b0 c0 d0 Col1: a1 b1 c1 d1 Col2: a2 b2 c2 d2 Col3: a3 b3 c3 d3
核心是通过vget_low_s32/vget_high_s32拆分向量为64位半部分,再用vzip1_s32/vzip2_s32交错取元素,最后用vcombine_s32重组为128位向量。
代码实现(C语言NEON Intrinsics)
#include <arm_neon.h> void transpose_4x4_int32(int32x4_t *dest, const int32x4_t *src) { // 输入:src[0]=Row0, src[1]=Row1, src[2]=Row2, src[3]=Row3 const int32x4_t row0 = src[0]; const int32x4_t row1 = src[1]; const int32x4_t row2 = src[2]; const int32x4_t row3 = src[3]; // 拆分每个行向量为高低64位(各含2个int32元素) int32x2_t r0_low = vget_low_s32(row0); // [a0, a1] int32x2_t r0_high = vget_high_s32(row0); // [a2, a3] int32x2_t r1_low = vget_low_s32(row1); // [b0, b1] int32x2_t r1_high = vget_high_s32(row1); // [b2, b3] int32x2_t r2_low = vget_low_s32(row2); // [c0, c1] int32x2_t r2_high = vget_high_s32(row2); // [c2, c3] int32x2_t r3_low = vget_low_s32(row3); // [d0, d1] int32x2_t r3_high = vget_high_s32(row3); // [d2, d3] // 交错重组得到转置后的列向量的高低半部分 int32x2_t col0_low = vzip1_s32(r0_low, r1_low); // [a0, b0] int32x2_t col0_high = vzip1_s32(r2_low, r3_low); // [c0, d0] int32x2_t col1_low = vzip2_s32(r0_low, r1_low); // [a1, b1] int32x2_t col1_high = vzip2_s32(r2_low, r3_low); // [c1, d1] int32x2_t col2_low = vzip1_s32(r0_high, r1_high); // [a2, b2] int32x2_t col2_high = vzip1_s32(r2_high, r3_high); // [c2, d2] int32x2_t col3_low = vzip2_s32(r0_high, r1_high); // [a3, b3] int32x2_t col3_high = vzip2_s32(r2_high, r3_high); // [c3, d3] // 组合高低半部分为完整的128位列向量 dest[0] = vcombine_s32(col0_low, col0_high); // [a0, b0, c0, d0] dest[1] = vcombine_s32(col1_low, col1_high); // [a1, b1, c1, d1] dest[2] = vcombine_s32(col2_low, col2_high); // [a2, b2, c2, d2] dest[3] = vcombine_s32(col3_low, col3_high); // [a3, b3, c3, d3] }
效率说明
- 所有操作均为NEON向量指令,无标量运算,保持了SIMD的高效性
- 总指令数固定,与输入数据无关,适合嵌入式ARM场景的性能需求
内容的提问来源于stack exchange,提问作者Bogi
相关产品推荐
相关产品推荐

