关于std::execution::par_unseq在蒙特卡洛Black-Scholes模拟中结果异常与性能问题的技术咨询
关于std::execution::par_unseq在蒙特卡洛Black-Scholes模拟中结果异常与性能问题的技术咨询
大家好,
之前因为忙犯了几个错误,已经修正了不同实现间的逻辑不一致问题,还把float溢出的地方改成了double。现在我更新了代码来测试蒙特卡洛Black-Scholes模拟的各种SIMD和并行选项,但遇到了两个很困惑的问题,想请教各位:
核心现象与疑问
结果一致性问题:
- 当模拟1亿个场景时,
std::execution::par_unseq版本计算出的现值(8.3)和其他4个版本(10.2-10.4)差异极大;par版本结果为~10.2,标量、C23 SIMD、AVX SIMD版本结果在~10.4左右。 - 但把场景数降到100万时,所有版本的结果几乎完全一致。
我想知道为什么par_unseq在大样本量下会出现这么离谱的偏差?
- 当模拟1亿个场景时,
性能问题:
par和par_unseq版本的执行时间居然比单线程标量版本还慢,即使修复bug后par_unseq比par略快,但还是不如标量和AVX SIMD版本。
明明用了并行执行,为什么反而更慢?
环境细节
- 编译器:g++ 15.1.0
- CPU配置:(8 X 2808 MHz CPU s)
CPU Caches: L1 Data 32 KiB (x4) L1 Instruction 32 KiB (x4) L2 Unified 256 KiB (x4) L3 Unified 6144 KiB (x1) - 系统负载:Load Average: 0.44, 0.52, 0.52
基准测试结果
1亿场景下的测试结果
------------------------------------------------------------------------------------------------------ Benchmark Time CPU Iterations UserCounters... ------------------------------------------------------------------------------------------------------ SimdFixture/black_scholes_mc_scalar 584217000 ns 584210200 ns 1 check_pv=10.4515 SimdFixture/black_scholes_mc_parallel 802403100 ns 802393700 ns 1 check_pv=10.2335 SimdFixture/black_scholes_mc_parallel_unseq 778954000 ns 778946000 ns 1 check_pv=8.322 SimdFixture/black_scholes_mc_simd_c23 601620500 ns 600948400 ns 1 check_pv=10.4515 SimdFixture/black_scholes_mc_simd_avx 466860950 ns 466861050 ns 2 check_pv=10.4191
100万场景下的测试结果
------------------------------------------------------------------------------------------------------ Benchmark Time CPU Iterations UserCounters... ------------------------------------------------------------------------------------------------------ SimdFixture/black_scholes_mc_scalar 7020675 ns 7020605 ns 110 check_pv=10.4656 SimdFixture/black_scholes_mc_parallel 6069825 ns 6069737 ns 110 check_pv=10.4657 SimdFixture/black_scholes_mc_parallel_unseq 6166214 ns 6166184 ns 112 check_pv=10.4655 SimdFixture/black_scholes_mc_simd_c23 6682331 ns 6682284 ns 99 check_pv=10.4656 SimdFixture/black_scholes_mc_simd_avx 4066819 ns 4066652 ns 149 check_pv=10.4657
实现差异说明
par和par_unseq版本:将每个场景的收益存入std::vector<float> payoff(N),再用std::reduce求和。- 其他版本:直接累积到标量double或SIMD向量float中,避免了大内存的分配和拷贝。
完整代码实现
#include "aligned_allocator.h" #include <algorithm> #include <benchmark/benchmark.h> #include <chrono> #include <cmath> #include <execution> // For C++17 parallel policies, or use simd with experimental::simd in C++23 #include <experimental/simd> #include <immintrin.h> // AVX intrinsics #include <iostream> #include <numeric> #include <random> #include <vector> typedef std::vector<float, aligned_allocator<float, 32>> AlignedVector; class SimdFixture : public benchmark::Fixture { protected: const int N = 100'000'000; std::vector<float, aligned_allocator<float, 32>> zs; float S0 = 100.0f, K = 100.0f, T = 1.0f, r = 0.05f, sigma = 0.2f; void SetUp(const ::benchmark::State &state) override { zs.resize(N); std::default_random_engine gen; std::normal_distribution<float> dist(0.0, 1.0); for (int i = 0; i < N; ++i) zs[i] = dist(gen); } }; BENCHMARK_F(SimdFixture, black_scholes_mc_scalar)(benchmark::State &state) { for (auto _ : state) { double pv = 0.0; double payoff_sum = 0.0f; float drift = (r - 0.5f * sigma * sigma) * T; float diffusion = sigma * std::sqrt(T); for (int i = 0; i < N; ++i) { float Z = zs[i]; float ST = S0 * std::exp(drift + diffusion * Z); payoff_sum += std::max(ST - K, 0.0f); } pv = std::exp(-r * T) * (payoff_sum / N); benchmark::DoNotOptimize(pv); state.counters["check_pv"] = pv; } } BENCHMARK_F(SimdFixture, black_scholes_mc_parallel)(benchmark::State &state) { for (auto _ : state) { double pv = 0.0; std::vector<float> payoff(N); float drift = (r - 0.5f * sigma * sigma) * T; float diffusion = sigma * std::sqrt(T); std::transform(std::execution::par, zs.begin(), zs.end(), payoff.begin(), [&](float z) { float ST = S0 * std::exp(drift + diffusion * z); return std::max(ST - K, 0.0f); }); double sum = std::reduce(std::execution::par, payoff.begin(), payoff.end()); pv = std::exp(-r * T) * (sum / N); benchmark::DoNotOptimize(pv); state.counters["check_pv"] = pv; } } BENCHMARK_F(SimdFixture, black_scholes_mc_parallel_unseq)(benchmark::State &state) { for (auto _ : state) { double pv = 0.0; std::vector<float> payoff(N); float drift = (r - 0.5f * sigma * sigma) * T; float diffusion = sigma * std::sqrt(T); std::transform(std::execution::par_unseq, zs.begin(), zs.end(), payoff.begin(), [&](float z) { float ST = S0 * std::exp(drift + diffusion * z); return std::max(ST - K, 0.0f); }); double sum = std::reduce(std::execution::par_unseq, payoff.begin(), payoff.end()); pv = std::exp(-r * T) * (sum / N); benchmark::DoNotOptimize(pv); state.counters["check_pv"] = pv; } } // 以下为省略的simd_c23和simd_avx实现(与用户提供的基准测试对应)
内容来源于stack exchange
相关产品推荐
相关产品推荐

