You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AVX优化版范围生成算法性能劣于朴素版本的问题排查及优化方案咨询

AVX优化版范围生成算法性能劣于朴素版本的问题排查及优化方案咨询

我在一个C语言项目中需要快速生成一个双精度浮点数的一维向量,只要给定起始值、步长和结束值就能生成。对我来说生成速度越快越好,所以一开始计划用SIMD intrinsics来做向量化优化。

我已经写出了基于AVX intrinsics的实现,功能上是正常的——比如调用generate_range_avx(0.0, 0.5, 3.0)时,能正确生成向量[0.0 0.5 1.0 1.5 2.0 2.5 3.0]。但实际测试下来,这个SIMD版本居然比朴素循环版本慢了2-3倍!

因为我对SIMD指令和intrinsics并不精通,所以我的向量化实现肯定存在问题,效率非常低下。下面是我的完整代码,想请教各位:

  • 我的代码到底哪里写得不对?
  • 这个实现还有优化的空间吗?
  • 是不是这类范围生成的问题本身就不适合用SIMD来高效优化?

我的开发环境是Windows/MSVC,目标架构为x86_64,因此使用了对应的编译器intrinsics。

#include <stdio.h>
#include <stdlib.h>
#include <memory.h>
#include <math.h>
#include <immintrin.h>

#define CALCELTS(start, increment, end) ((size_t)round(((end) - (start)) / (increment)) + 1)

static size_t align(size_t num, size_t alignment) {
    return (num + alignment - 1) & ~(alignment - 1);
}

typedef struct {
    double* elts;
    size_t len;
} vector_t;

vector_t generate_range_naive(double start, double increment, double end) {
    // Calculate the unaligned number of elements
    size_t num_elements = CALCELTS(start, increment, end);
    size_t aligned_num_elts = align(num_elements, 4);

    double* data = malloc(sizeof * data * aligned_num_elts);
    if (!data) {
        fprintf(stderr, "Failed to allocate memory.");
        exit(1);
    }

    for (size_t i = 0; i < num_elements; i++) {
        data[i] = start;
        start += increment;
    }
    data[num_elements - 1] = end;

    return (vector_t) { data, num_elements };
}

vector_t generate_range_avx(double start, double increment, double end) {
    // Calculate the unaligned number of elements
    size_t num_elements = CALCELTS(start, increment, end);
    if (num_elements < 4)
        return generate_range_naive(start, increment, end);

    size_t aligned_num_elts = align(num_elements, 4);
    // Allocate memory for the vector.
    double* data = _mm_malloc(sizeof(*data) * aligned_num_elts, 32);
    if (!data) {
        fprintf(stderr, "Failed to allocate memory.");
        exit(1);
    }

    __m256d start_v = _mm256_set1_pd(start);
    __m256d indices_v = _mm256_set_pd(3.0, 2.0, 1.0, 0.0);
    __m256d increment_v = _mm256_set1_pd(increment);
    __m256d x_v = _mm256_add_pd(start_v, _mm256_mul_pd(indices_v, increment_v));

    increment_v = _mm256_mul_pd(increment_v, _mm256_set1_pd(4.0));

    for (size_t i = 0; i < aligned_num_elts; i += 4) {
        _mm256_stream_pd(data + i, x_v);
        x_v = _mm256_add_pd(x_v, increment_v);
    }

    data[num_elements - 1] = end;
    return (vector_t) { data, num_elements };
}

static void print_range(vector_t vec) {
    for (size_t i = 0; i < vec.len; i++) {
        printf("%lf ", vec.elts[i]);
    }
    putchar('\n');
}

int main(int argc, char* argv[]) {
    // 我用自己的简单基准测试库做了性能测试,这里就没贴代码了,如果需要的话可以补充
    vector_t vec1 = generate_range_avx(1.0, 0.5, 100.0);
    print_range(vec1);

    vector_t vec2 = generate_range_naive(1.0, 0.5, 100.0);
    print_range(vec2);

    return 0;
}

备注:内容来源于stack exchange,提问作者jtxkopt - STOP GENOCIDE

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.20 11:30:29