You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

C++函数各模块执行时间总和与整体耗时不符问题求助

程序执行时间异常翻倍问题排查

我正在优化程序执行时间,现有KAT类包含references、samples、pointerMatrix等成员变量,当前有3个reference和3个sample。由于refAligner.align(sample)返回结果体积较大,我使用了std::unique_ptr智能指针。单独测量各部分执行时间总和在400-500秒,但整个alignAllSamplesToAllReferences函数执行时间却是这个数值的两倍。注意到最后一条"Add to matrix:"的std::cout输出与"Entire alignment:"的输出之间有很长等待时间,想知道额外耗时的来源,可提供更多代码或信息。


KAT类定义

class KAT {

    private:
        std::string REFERENCES_FOLDER = "/References";
        std::string SAMPLES_FOLDER = "/Samples";
        Input input;
        References references;
        Filenames referenceFilenames;
        Samples samples;
        Filenames sampleFilenames;
        std::vector<std::vector<std::unique_ptr<Alignment>>> pointerMatrix;
        Results result;

        void readReferences();
        void calculateReferencesTextIndex();
        void readSamples();
        void createKmers();
        void alignAllSamplesToAllReferences();
        Results createResult();
        void copyAlignmentsToResult();


    public:
        KAT(Input input_): input(input_) { }
        
        Results align();

};

问题方法:alignAllSamplesToAllReferences

void KAT::alignAllSamplesToAllReferences() {

    auto start1 = std::chrono::high_resolution_clock::now();
    
    for (Reference reference : references) {

        auto start2 = std::chrono::high_resolution_clock::now();
        
        ReferenceAligner refAligner(reference);
        std::vector<std::unique_ptr<Alignment>> alignPointers;
        
        auto stop2 = std::chrono::high_resolution_clock::now();
        auto duration2 = std::chrono::duration_cast<std::chrono::seconds>(stop2 - start2);
        std::cout << "RefAligner Class: " << duration2.count() << std::endl;

        for (Sample sample: samples) {

            auto start3 = std::chrono::high_resolution_clock::now();

            std::unique_ptr<Alignment> aligmentPointer = refAligner.align(sample);
            alignPointers.push_back(std::move(aligmentPointer));
            
            auto stop3 = std::chrono::high_resolution_clock::now();
            auto duration3 = std::chrono::duration_cast<std::chrono::seconds>(stop3 - start3);
            std::cout << "Align and pushback: " << duration3.count() << std::endl;
        }

        auto start4 = std::chrono::high_resolution_clock::now();

        pointerMatrix.push_back(std::move(alignPointers));

        auto stop4 = std::chrono::high_resolution_clock::now();
        auto duration4 = std::chrono::duration_cast<std::chrono::seconds>(stop4 - start4);
        std::cout << "Add to matrix: " << duration4.count() << std::endl;
    }

    auto stop1 = std::chrono::high_resolution_clock::now();
    auto duration1 = std::chrono::duration_cast<std::chrono::seconds>(stop1 - start1);
    std::cout << "Entire alignment: " << duration1.count() << std::endl;
}

执行时间结果

Aligning...
RefAligner Class: 0
Align and pushback: 12
Align and pushback: 44
Align and pushback: 21
Add to matrix: 0
RefAligner Class: 0
Align and pushback: 18
Align and pushback: 70
Align and pushback: 35
Add to matrix: 0
RefAligner Class: 0
Align and pushback: 46
Align and pushback: 115
Align and pushback: 53
Add to matrix: 0
Entire alignment: 999
Align all samples: 999
Create Results: 0
Writing results...
Main: 1171

额外耗时的可能原因及解决方向

  1. 循环中的对象拷贝开销
    当前循环使用值传递遍历references和samples:

    for (Reference reference : references)
    for (Sample sample: samples)
    

    如果Reference或Sample是大体积对象,每次迭代都会触发完整拷贝,这部分开销完全没被你的分段计时覆盖。改成const引用传递就能避免不必要的拷贝:

    for (const Reference& reference : references)
    for (const Sample& sample : samples)
    
  2. std::cout的同步与IO延迟
    循环内频繁调用std::cout输出,默认情况下std::cout会和C标准IO同步,每次输出都有同步开销,且实际输出可能延迟到后续执行阶段,这部分耗时没被计入分段计时。可以在函数开头添加以下代码关闭同步:

    std::ios_base::sync_with_stdio(false);
    std::cin.tie(nullptr);
    

    或者将所有输出内容缓存到std::stringstream中,最后一次性输出,减少IO操作次数。

  3. 对象析构的隐藏开销
    每个循环迭代中创建的ReferenceAligner对象,以及alignPointers中管理的Alignment对象,会在循环迭代结束后被析构。如果ReferenceAligner或Alignment的析构函数包含大量清理工作(比如释放大块内存、关闭资源),这部分耗时发生在Add to matrix:输出之后、Entire alignment:计时之前,完全不在你的分段计时范围内。可以给这两个类的析构函数添加计时,确认是否是这部分导致的延迟。

  4. vector扩容的隐性开销
    pointerMatrix在push_back时如果容量不足,会触发内存重新分配和元素移动。虽然你的Add to matrix:计时显示为0秒,但多次扩容的累计开销可能被忽略。可以在循环开始前给pointerMatrix预留足够容量:

    pointerMatrix.reserve(references.size());
    
  5. 编译优化等级问题
    如果当前是Debug模式编译,会有大量调试相关的额外开销,导致执行时间大幅增加。确保编译时开启合适的优化等级(比如-O2或-O3),再测试执行时间。


内容的提问来源于stack exchange,提问作者efrain_ceh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 01:17:05