同优化配置下Linux编译C++程序比Windows慢的原因与优化问询
更新2
如有需要,可通过GitHub获取原始代码,也可在编辑历史中找到复现问题所做的完整精确修改及程序日志(这些细节会使问题偏离重点)。
更新并再次强调:此问题与C++程序计时方法无关
正如原始问题所述,我专门测量了实际耗时(挂钟时间),Windows下为20秒,Linux下为60秒。我用手机秒表确认了这一结果。我的唯一问题是:为何启用相同优化特性的程序在Linux上比Windows慢这么多?
我尝试在Linux上运行某GitHub代码,发现其运行速度比Windows慢2到3倍。使用官方输入示例data/BS_1000_torus.xyz,Windows下耗时约20秒,Linux下耗时约60秒**(我用手机秒表确认了这一结果)**。我正尝试调整编译配置,使Linux下的运行性能与Windows持平,以下是详细说明:
Windows环境下:
我严格遵循README中的步骤(启用AVX2、快速浮点运算及OpenMP),使用vcpkg和VS2022编译项目。运行BS_1000_torus耗时约20秒(经秒表确认)。
Linux环境下:
在Linux上,我对CMakeLists.txt做了以下修改以启用README中提及的关键特性:
- 移除
CMakeLists.txt中的vcpkg工具链指定 - 在
set(CMAKE_BUILD_TYPE RELEASE)后添加set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mavx2 -fopenmp -pthread -Ofast"),以启用README中提到的特性
随后我使用cmake和make编译:
mkdir build cd build cmake .. make -j
cmake输出如下:
-- The C compiler identification is GNU 9.4.0 -- The CXX compiler identification is GNU 9.4.0 -- Detecting C compiler ABI info -- Detecting C compiler ABI info - done -- Check for working C compiler: /usr/bin/cc - skipped -- Detecting C compile features -- Detecting C compile features - done -- Detecting CXX compiler ABI info -- Detecting CXX compiler ABI info - done -- Check for working CXX compiler: /usr/bin/c++ - skipped -- Detecting CXX compile features -- Detecting CXX compile features - done -- 3.3.9 -- Found Boost: /usr/lib/x86_64-linux-gnu/cmake/Boost-1.71.0/BoostConfig.cmake (found version "1.71.0") -- BOOST FOUNDED -- Using header-only CGAL -- Targeting Unix Makefiles -- Using /usr/bin/c++ compiler. -- Found GMP: /usr/lib/x86_64-linux-gnu/libgmp.so -- Found MPFR: /usr/lib/x86_64-linux-gnu/libmpfr.so -- Found Boost: /usr/lib/x86_64-linux-gnu/cmake/Boost-1.71.0/BoostConfig.cmake (found suitable version "1.71.0", minimum required is "1.66") -- Boost include dirs: /usr/include -- Boost libraries: -- Performing Test CMAKE_HAVE_LIBC_PTHREAD -- Performing Test CMAKE_HAVE_LIBC_PTHREAD - Failed -- Check if compiler accepts -pthread -- Check if compiler accepts -pthread - yes -- Found Threads: TRUE -- Using gcc version 4 or later. Adding -frounding-math -- Build type: RELEASE -- USING CXXFLAGS = ' -mavx2 -fopenmp -pthread -Ofast -O3 -DNDEBUG' -- USING EXEFLAGS = ' ' -- Found OpenMP_C: -fopenmp (found version "4.5") -- Found OpenMP_CXX: -fopenmp (found version "4.5") -- Found OpenMP: TRUE (found version "4.5") -- Configuring done (1.2s) -- Generating done (0.0s) -- Build files have been written to: /home/user/3dlab/GCNO-master/build
运行与Windows实验相同的模型耗时约64秒(经秒表确认)。
系统配置:
- 两次测试在同一台电脑(双系统启动,非WSL)上进行,CPU为Intel(R) Core(TM) i9-10900X @ 3.70GHz(10核,每核2线程)。我在
int main开头设置了omp_set_num_threads(20); - Windows系统:Windows 10
- Linux系统:5.15.0-88-generic #98~20.04.1-Ubuntu
问题:
- 为何即使在Linux上启用了所有能想到的优化标志,运行时间(实际耗时)仍差异巨大(Windows 20秒 vs Linux 60秒)?为何启用相同特性(AVX2、OpenMP)编译后运行时间差异如此之大?
- 如何调整编译配置使Linux下的运行速度与Windows持平?是否存在Windows自动启用但Linux需手动开启的优化项?
内容的提问来源于stack exchange,提问作者ihdv
相关产品推荐
相关产品推荐

