You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

并行化大复数短向量填充时的内存溢出问题排查

问题描述

需要处理一个容量为72亿的std::vector<std::complex<short>>类型向量,用于存储复数短型数据。单线程下预分配内存后循环填充需耗时4分钟,占用约80-100GB内存。

尝试用OpenMP并行化填充逻辑,为循环添加#pragma omp parallel for,并通过#pragma omp critical保护push_back操作,但64线程满负载运行后内存瞬间耗尽至251GB,程序被Linux以SIGKILL信号终止。

附可复现代码:

#include <iostream>
#include <vector>
#include <complex>
#include <math.h>
#include <signal.h>
#include <iostream>
#include "/usr/lib/llvm-14/lib/clang/14.0.0/include/omp.h"

int main()
{
        // this vector holds the total samples that define waveform
        std::vector<std::complex<short>> pulseSequence;

        float duration = 180; // 180 seconds chirp sounding

        float sampleRate = 40e6; // sample beyond Nyquist limit + prevent aliasing on SDR

        size_t nSamples = duration * sampleRate;

        std::complex<short> sample; // store wavepacket

        pulseSequence.reserve(nSamples); //allocate memory for large vector (7,200,000,000 samples)

        #pragma omp parallel for
        for ( size_t iSample=0; iSample < nSamples; iSample++ )
        {
                // the following is bogus/test code for example purposes:
                sample = std::complex<short>(0*iSample,0);

                //save the sample into the main vector
                #pragma omp critical // <-- tried with/without
                {
                        pulseSequence.push_back(sample);
                }

                //output some progress:
                if (iSample % 100000000 == 0)
                {
                    std::cout << float(iSample)*100.0/float(nSamples) << "%" << std::endl;
                }

        }

        return 0; // end program
}

编译命令:g++ test.cpp -o test -fopenmp

内存耗尽的原因与错误分析
  • vector使用逻辑错误:你调用reserve(nSamples)仅预留了内存容量,但vector的实际元素数量(size)仍为0。此时用push_back逐个添加元素,在并行场景下:
    • 不加critical时,多线程并发执行push_back会直接破坏vector内部的size计数器、尾指针等结构,导致vector错误触发额外内存分配甚至堆损坏,最终内存暴涨。
    • 加critical后,所有线程串行执行push_back,完全失去并行加速意义,且vector频繁串行修改时的内部开销可能引发内存异常。
  • 共享变量数据冲突:std::complex<short> sample定义在并行循环外,属于线程共享变量,多线程同时修改会导致写入vector的样本数据被覆盖,同时可能产生额外的内存同步开销。
  • 并行输出冗余开销:每个线程都会执行进度输出逻辑,大量并发的cout操作会阻塞IO,缓冲区累积也会占用额外内存。
修复方案

核心思路是让线程直接对vector的对应索引位置赋值,彻底规避push_back的线程安全问题,最大化并行效率:

#include <iostream>
#include <vector>
#include <complex>
#include <math.h>
#include <signal.h>
#include <omp.h>

int main()
{
        std::vector<std::complex<short>> pulseSequence;

        float duration = 180;
        float sampleRate = 40e6;
        size_t nSamples = duration * sampleRate;

        // 用resize初始化元素数量与内存,替代reserve
        pulseSequence.resize(nSamples);

        #pragma omp parallel for
        for (size_t iSample = 0; iSample < nSamples; iSample++)
        {
                // 线程私有变量,避免数据覆盖
                std::complex<short> sample(0 * iSample, 0);
                // 直接对索引位置赋值,无需同步保护
                pulseSequence[iSample] = sample;

                // 仅主线程输出进度,减少IO冲突
                if (iSample % 100000000 == 0 && omp_get_thread_num() == 0)
                {
                        std::cout << static_cast<float>(iSample) * 100.0 / static_cast<float>(nSamples) << "%" << std::endl;
                }
        }

        return 0;
}

编译命令保持不变:g++ test.cpp -o test -fopenmp

额外说明

修复后每个线程独立处理自己的迭代区间,直接对vector的对应位置赋值,无需任何同步操作,并行效率最大化;内存占用与单线程一致(约80-100GB),无额外冗余分配。

内容的提问来源于stack exchange,提问作者NoRainDropsInTheSky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.14 10:10:01