You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Zen4架构下__m512i中64位值成对相加的实现方案咨询

在AMD Zen4平台上的实现方案

Zen4完全支持AVX-512指令集,针对你提出的三种结果,以下是具体的实现方法:

1. 得到包含4个128位值的__m512i向量

_mm512_popcnt_epi64()输出的是8个64位整数组成的__m512i向量,成对相加(第0和1、2和3、4和5、6和7位相加)并保留128位结果,可通过排列向量后扩展相加实现:

__m512i popcnt_vec = _mm512_popcnt_epi64(input_vec);
// 生成索引,将奇数位元素移到对应偶数位的下一个位置
__m512i indices = _mm512_set_epi64(7,5,3,1,6,4,2,0);
__m512i shifted = _mm512_permutexvar_epi64(indices, popcnt_vec);

// 将原向量和移位后的向量扩展为128位整数,再成对相加
__m512i extended_src = _mm512_cvtepi64_epi128(popcnt_vec);
__m512i extended_shifted = _mm512_cvtepi64_epi128(shifted);
__m512i final_result = _mm512_add_epi128(extended_src, extended_shifted);

2. 得到包含4个64位值的__m256i向量

直接对成对的64位值相加,再提取结果压缩到256位向量:

__m512i popcnt_vec = _mm512_popcnt_epi64(input_vec);
// 移位后让成对元素对齐
__m512i indices = _mm512_set_epi64(7,5,3,1,6,4,2,0);
__m512i shifted = _mm512_permutexvar_epi64(indices, popcnt_vec);

// 成对相加后,提取偶数位的4个结果到256位向量
__m512i sum = _mm512_add_epi64(popcnt_vec, shifted);
__m256i final_result = _mm512_extracti64x4_epi64(sum, 0);

3. 得到包含4个32位值的__m128i向量

由于64位popcount的结果最大为64,可安全转换为32位,再成对相加并压缩:

__m512i popcnt_vec = _mm512_popcnt_epi64(input_vec);
// 将64位结果转换为32位
__m512i popcnt_32 = _mm512_cvtepi64_epi32(popcnt_vec);

// 生成索引让成对的32位元素对齐,相加后提取结果
__m512i indices = _mm512_set_epi32(15,13,11,9,7,5,3,1,14,12,10,8,6,4,2,0);
__m512i shifted = _mm512_permutexvar_epi32(indices, popcnt_32);
__m512i sum_32 = _mm512_add_epi32(popcnt_32, shifted);

__m128i final_result = _mm512_extracti32x4_epi32(sum_32, 0);

所有实现均充分利用Zen4的AVX-512指令集特性,减少冗余数据移动,保证执行效率。

内容的提问来源于stack exchange,提问作者John Smith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 10:20:26