You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于std::execution::par_unseq在蒙特卡洛Black-Scholes模拟中结果异常与性能问题的技术咨询

关于std::execution::par_unseq在蒙特卡洛Black-Scholes模拟中结果异常与性能问题的技术咨询

大家好,

之前因为忙犯了几个错误,已经修正了不同实现间的逻辑不一致问题,还把float溢出的地方改成了double。现在我更新了代码来测试蒙特卡洛Black-Scholes模拟的各种SIMD和并行选项,但遇到了两个很困惑的问题,想请教各位:

核心现象与疑问

  1. 结果一致性问题:

    • 当模拟1亿个场景时,std::execution::par_unseq版本计算出的现值(8.3)和其他4个版本(10.2-10.4)差异极大;par版本结果为~10.2,标量、C23 SIMD、AVX SIMD版本结果在~10.4左右。
    • 但把场景数降到100万时,所有版本的结果几乎完全一致。
      我想知道为什么par_unseq在大样本量下会出现这么离谱的偏差?
  2. 性能问题:

    • par和par_unseq版本的执行时间居然比单线程标量版本还慢,即使修复bug后par_unseq比par略快,但还是不如标量和AVX SIMD版本。
      明明用了并行执行,为什么反而更慢?

环境细节

  • 编译器:g++ 15.1.0
  • CPU配置:(8 X 2808 MHz CPU s)
    CPU Caches:
      L1 Data 32 KiB (x4)
      L1 Instruction 32 KiB (x4)
      L2 Unified 256 KiB (x4)
      L3 Unified 6144 KiB (x1)
    
  • 系统负载:Load Average: 0.44, 0.52, 0.52

基准测试结果

1亿场景下的测试结果

------------------------------------------------------------------------------------------------------
Benchmark                                  Time             CPU   Iterations UserCounters...
------------------------------------------------------------------------------------------------------
SimdFixture/black_scholes_mc_scalar     584217000 ns    584210200 ns            1 check_pv=10.4515
SimdFixture/black_scholes_mc_parallel   802403100 ns    802393700 ns            1 check_pv=10.2335
SimdFixture/black_scholes_mc_parallel_unseq 778954000 ns 778946000 ns            1 check_pv=8.322
SimdFixture/black_scholes_mc_simd_c23   601620500 ns    600948400 ns            1 check_pv=10.4515
SimdFixture/black_scholes_mc_simd_avx   466860950 ns    466861050 ns            2 check_pv=10.4191

100万场景下的测试结果

------------------------------------------------------------------------------------------------------
Benchmark                                  Time             CPU   Iterations UserCounters...
------------------------------------------------------------------------------------------------------
SimdFixture/black_scholes_mc_scalar       7020675 ns       7020605 ns          110 check_pv=10.4656
SimdFixture/black_scholes_mc_parallel     6069825 ns       6069737 ns          110 check_pv=10.4657
SimdFixture/black_scholes_mc_parallel_unseq 6166214 ns    6166184 ns          112 check_pv=10.4655
SimdFixture/black_scholes_mc_simd_c23     6682331 ns       6682284 ns           99 check_pv=10.4656
SimdFixture/black_scholes_mc_simd_avx     4066819 ns       4066652 ns          149 check_pv=10.4657

实现差异说明

  • par和par_unseq版本:将每个场景的收益存入std::vector<float> payoff(N),再用std::reduce求和。
  • 其他版本:直接累积到标量double或SIMD向量float中,避免了大内存的分配和拷贝。

完整代码实现

#include "aligned_allocator.h"
#include <algorithm>
#include <benchmark/benchmark.h>
#include <chrono>
#include <cmath>
#include <execution> // For C++17 parallel policies, or use simd with experimental::simd in C++23
#include <experimental/simd>
#include <immintrin.h> // AVX intrinsics
#include <iostream>
#include <numeric>
#include <random>
#include <vector>

typedef std::vector<float, aligned_allocator<float, 32>> AlignedVector;

class SimdFixture : public benchmark::Fixture {
protected:
    const int N = 100'000'000;
    std::vector<float, aligned_allocator<float, 32>> zs;
    float S0 = 100.0f, K = 100.0f, T = 1.0f, r = 0.05f, sigma = 0.2f;

    void SetUp(const ::benchmark::State &state) override {
        zs.resize(N);
        std::default_random_engine gen;
        std::normal_distribution<float> dist(0.0, 1.0);
        for (int i = 0; i < N; ++i)
            zs[i] = dist(gen);
    }
};

BENCHMARK_F(SimdFixture, black_scholes_mc_scalar)(benchmark::State &state) {
    for (auto _ : state) {
        double pv = 0.0;
        double payoff_sum = 0.0f;
        float drift = (r - 0.5f * sigma * sigma) * T;
        float diffusion = sigma * std::sqrt(T);
        for (int i = 0; i < N; ++i) {
            float Z = zs[i];
            float ST = S0 * std::exp(drift + diffusion * Z);
            payoff_sum += std::max(ST - K, 0.0f);
        }
        pv = std::exp(-r * T) * (payoff_sum / N);
        benchmark::DoNotOptimize(pv);
        state.counters["check_pv"] = pv;
    }
}

BENCHMARK_F(SimdFixture, black_scholes_mc_parallel)(benchmark::State &state) {
    for (auto _ : state) {
        double pv = 0.0;
        std::vector<float> payoff(N);
        float drift = (r - 0.5f * sigma * sigma) * T;
        float diffusion = sigma * std::sqrt(T);
        std::transform(std::execution::par, zs.begin(), zs.end(), payoff.begin(),
                       [&](float z) {
                           float ST = S0 * std::exp(drift + diffusion * z);
                           return std::max(ST - K, 0.0f);
                       });
        double sum = std::reduce(std::execution::par, payoff.begin(), payoff.end());
        pv = std::exp(-r * T) * (sum / N);
        benchmark::DoNotOptimize(pv);
        state.counters["check_pv"] = pv;
    }
}

BENCHMARK_F(SimdFixture, black_scholes_mc_parallel_unseq)(benchmark::State &state) {
    for (auto _ : state) {
        double pv = 0.0;
        std::vector<float> payoff(N);
        float drift = (r - 0.5f * sigma * sigma) * T;
        float diffusion = sigma * std::sqrt(T);
        std::transform(std::execution::par_unseq, zs.begin(), zs.end(), payoff.begin(),
                       [&](float z) {
                           float ST = S0 * std::exp(drift + diffusion * z);
                           return std::max(ST - K, 0.0f);
                       });
        double sum = std::reduce(std::execution::par_unseq, payoff.begin(), payoff.end());
        pv = std::exp(-r * T) * (sum / N);
        benchmark::DoNotOptimize(pv);
        state.counters["check_pv"] = pv;
    }
}

// 以下为省略的simd_c23和simd_avx实现(与用户提供的基准测试对应)

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.07 10:48:08