C++函数各模块执行时间总和与整体耗时不符问题求助
我正在优化程序执行时间,现有KAT类包含references、samples、pointerMatrix等成员变量,当前有3个reference和3个sample。由于refAligner.align(sample)返回结果体积较大,我使用了std::unique_ptr智能指针。单独测量各部分执行时间总和在400-500秒,但整个alignAllSamplesToAllReferences函数执行时间却是这个数值的两倍。注意到最后一条"Add to matrix:"的std::cout输出与"Entire alignment:"的输出之间有很长等待时间,想知道额外耗时的来源,可提供更多代码或信息。
KAT类定义
class KAT { private: std::string REFERENCES_FOLDER = "/References"; std::string SAMPLES_FOLDER = "/Samples"; Input input; References references; Filenames referenceFilenames; Samples samples; Filenames sampleFilenames; std::vector<std::vector<std::unique_ptr<Alignment>>> pointerMatrix; Results result; void readReferences(); void calculateReferencesTextIndex(); void readSamples(); void createKmers(); void alignAllSamplesToAllReferences(); Results createResult(); void copyAlignmentsToResult(); public: KAT(Input input_): input(input_) { } Results align(); };
问题方法:alignAllSamplesToAllReferences
void KAT::alignAllSamplesToAllReferences() { auto start1 = std::chrono::high_resolution_clock::now(); for (Reference reference : references) { auto start2 = std::chrono::high_resolution_clock::now(); ReferenceAligner refAligner(reference); std::vector<std::unique_ptr<Alignment>> alignPointers; auto stop2 = std::chrono::high_resolution_clock::now(); auto duration2 = std::chrono::duration_cast<std::chrono::seconds>(stop2 - start2); std::cout << "RefAligner Class: " << duration2.count() << std::endl; for (Sample sample: samples) { auto start3 = std::chrono::high_resolution_clock::now(); std::unique_ptr<Alignment> aligmentPointer = refAligner.align(sample); alignPointers.push_back(std::move(aligmentPointer)); auto stop3 = std::chrono::high_resolution_clock::now(); auto duration3 = std::chrono::duration_cast<std::chrono::seconds>(stop3 - start3); std::cout << "Align and pushback: " << duration3.count() << std::endl; } auto start4 = std::chrono::high_resolution_clock::now(); pointerMatrix.push_back(std::move(alignPointers)); auto stop4 = std::chrono::high_resolution_clock::now(); auto duration4 = std::chrono::duration_cast<std::chrono::seconds>(stop4 - start4); std::cout << "Add to matrix: " << duration4.count() << std::endl; } auto stop1 = std::chrono::high_resolution_clock::now(); auto duration1 = std::chrono::duration_cast<std::chrono::seconds>(stop1 - start1); std::cout << "Entire alignment: " << duration1.count() << std::endl; }
执行时间结果
Aligning... RefAligner Class: 0 Align and pushback: 12 Align and pushback: 44 Align and pushback: 21 Add to matrix: 0 RefAligner Class: 0 Align and pushback: 18 Align and pushback: 70 Align and pushback: 35 Add to matrix: 0 RefAligner Class: 0 Align and pushback: 46 Align and pushback: 115 Align and pushback: 53 Add to matrix: 0 Entire alignment: 999 Align all samples: 999 Create Results: 0 Writing results... Main: 1171
额外耗时的可能原因及解决方向
循环中的对象拷贝开销
当前循环使用值传递遍历references和samples:for (Reference reference : references) for (Sample sample: samples)如果
Reference或Sample是大体积对象,每次迭代都会触发完整拷贝,这部分开销完全没被你的分段计时覆盖。改成const引用传递就能避免不必要的拷贝:for (const Reference& reference : references) for (const Sample& sample : samples)std::cout的同步与IO延迟
循环内频繁调用std::cout输出,默认情况下std::cout会和C标准IO同步,每次输出都有同步开销,且实际输出可能延迟到后续执行阶段,这部分耗时没被计入分段计时。可以在函数开头添加以下代码关闭同步:std::ios_base::sync_with_stdio(false); std::cin.tie(nullptr);或者将所有输出内容缓存到
std::stringstream中,最后一次性输出,减少IO操作次数。对象析构的隐藏开销
每个循环迭代中创建的ReferenceAligner对象,以及alignPointers中管理的Alignment对象,会在循环迭代结束后被析构。如果ReferenceAligner或Alignment的析构函数包含大量清理工作(比如释放大块内存、关闭资源),这部分耗时发生在Add to matrix:输出之后、Entire alignment:计时之前,完全不在你的分段计时范围内。可以给这两个类的析构函数添加计时,确认是否是这部分导致的延迟。vector扩容的隐性开销
pointerMatrix在push_back时如果容量不足,会触发内存重新分配和元素移动。虽然你的Add to matrix:计时显示为0秒,但多次扩容的累计开销可能被忽略。可以在循环开始前给pointerMatrix预留足够容量:pointerMatrix.reserve(references.size());编译优化等级问题
如果当前是Debug模式编译,会有大量调试相关的额外开销,导致执行时间大幅增加。确保编译时开启合适的优化等级(比如-O2或-O3),再测试执行时间。
内容的提问来源于stack exchange,提问作者efrain_ceh

