Cython封装的C++ Vector函数性能远低于纯Cython NumPy MemoryView函数的优化方案问询
我现在遇到一个性能优化的问题,想请教大家怎么改进Cython封装C++函数的性能。情况是这样的:
我分别用两种方式实现了逻辑完全等价的算法:一种是C++编写的cpp_func1(通过Cython封装供Python调用),处理的是std::vector容器;另一种是纯Cython编写的cython_func1,基于NumPy数组和MemoryView实现。当我用包含10万个元素的NumPy数组从Python调用这两个函数时,发现封装后的cpp_func1执行速度至少比纯Cython版本慢4倍,内存占用更是高了3倍多。
更夸张的是,在另一个逻辑类似的测试案例(cpp_func2和cython_func2)里,封装的C函数平均执行速度比Cython版本慢了整整24倍。目前纯Cython的cython_func1性能表现最好,不过还没尝试并行优化;但我主要想解决的是C函数封装后的性能问题,肯定是我漏掉了关键的优化点,希望大家能给点针对性的建议。
做这个基准测试的核心目的,是对比循环+容器的等价实现,尽量通过容器引用避免数据拷贝,以此保证处理大型容器时的高效性。
纯Cython实现代码
cython_funcs.pyx
def cython_func1(double[::1] arr): cdef int n, m cdef int window = 5 cdef int len = arr.shape[0] - window + 1 cdef double min_var = 0.0 cdef double max_var = 0.0 cdef double diff_var = 0.0 arr_result = np.zeros(len, dtype=np.double) cdef double[::1] arr_result_view = arr_result for n in range(len): diff_var = 0.0 for m in range(n, (n + window)): if (m == n): min_var = arr[m] max_var = arr[m] else: if (arr[m] < min_var): min_var = arr[m] elif (arr[m] > max_var): max_var = arr[m] diff_var = max_var - min_var arr_result_view[n] = diff_var return arr_result
cython_funcs_setup.py
# 编译命令(CMD): # python cython_funcs_setup.py build_ext --inplace from setuptools import setup from Cython.Build import cythonize import numpy setup( ext_modules = cythonize("cython_funcs.pyx"), include_path = [numpy.get_include()] )
C++函数+Cython封装代码
cpp_funcs.cpp
std::vector<double> cpp_func1(const std::vector<double> &vec) { int window = 5; int len = vec.size() - window + 1; double min_var = 0.0; double max_var = 0.0; double diff_var = 0.0; std::vector<double> vec_result(len, 0.0); for (int n = 0; n < len; n++) { diff_var = 0.0; for (int m = n; m < (n + window); m++) { if (m == n) { min_var = vec[m]; max_var = vec[m]; } else { if (vec[m] < min_var) { min_var = vec[m]; } else if (vec[m] > max_var) { max_var = vec[m]; } } } diff_var = max_var - min_var; vec_result[n] = diff_var; } return vec_result; }
cpp_funcs.h
#ifndef CPP_FUNCS_H #define CPP_FUNCS_H #include <vector> #include <stdint.h> using namespace std; std::vector<double> cpp_func1(const std::vector<double> &vec); #endif
cpp_funcs_wrapper.pyx
import numpy as np from libcpp.vector cimport vector #cdef extern from "cpp_funcs.cpp": # pass cdef extern from "cpp_funcs.h": vector[double] cpp_func1(vector[double] &vec) def cpp_func1_cython(vector[double] &vec): return cpp_func1(vec)
cpp_funcs_setup.py
# 编译命令(CMD): # python cpp_funcs_setup.py build_ext --inplace from setuptools import setup from distutils.extension import Extension from Cython.Build import cythonize import numpy extensions = [Extension("cpp_funcs_wrapper", ["cpp_funcs_wrapper.pyx", "cpp_funcs.cpp"], language="c++", #extra_link_args=["-lz"] )] setup( name="cpp_funcs_wrapper", ext_modules=cythonize(extensions), include_path = [numpy.get_include()] )
Python调用测试代码
import cython_funcs import cpp_funcs_wrapper import numpy as np import timeit arr = np.random.uniform(1, 100000, 100_000) var1 = timeit.timeit("cython_funcs.cython_func1(arr)", globals=globals(), number=100) print(f"Average time = {((var1/100) * 1000)} ms.") var2 = timeit.timeit("cpp_funcs_wrapper.cpp_func1_cython(arr)", globals=globals(), number=100) print(f"Average time = {((var2/100) * 1000)} ms")
已尝试的优化及问题(更新1-3)
更新1:尝试编译优化但未生效
我试过在Cython的setup文件里添加编译优化标志(比如/O2),但要么标志被终端提示"unrecognized option '/O2'; ignored"而忽略,要么直接出现编译错误。即使编译成功,函数性能也没有任何变化。我前后试了十几种标志组合,举个典型的尝试案例:
cpp_funcs_O_setup.py
''' 以下不是当前使用的Cython setup文件,仅用于展示我试过的无效优化尝试。 执行时间测试并非基于该文件编译的结果,因为它完全没起到优化作用。 参考语法来源:Cython并行化文档的编译部分 所有优化尝试都只修改了Cython setup文件,编译命令未做改动。 编译命令(CMD): python cpp_funcs_O_setup.py build_ext --inplace ''' from setuptools import Extension, setup from Cython.Build import cythonize import sys # 需要时也添加了NumPy的头文件路径 if sys.platform.startswith("win"): compile_args = '/O2' else: compile_args = '-O2' ext_modules = [ Extension( "cpp_funcs_wrapper", sources=["cpp_funcs_wrapper.pyx", "cpp_funcs.cpp"], language="c++", extra_compile_args=[compile_args], # 也直接试过'/O1', '/O3', '/Ofast' '-O1'等多种标志 extra_link_args=[compile_args], ), # 也试过注释掉下面这个Extension,或者其他变体组合 Extension( "cpp_funcs_wrapper", sources=["cpp_funcs_wrapper.pyx", "cpp_funcs.cpp"], language="c++", extra_compile_args=[compile_args], extra_link_args=[compile_args], ) ] setup( name='cpp_funcs_wrapper', ext_modules=cythonize(ext_modules), )
更新2:编译时Visual Studio自动调用的情况
后来我注意到,执行编译命令python cpp_funcs_setup.py build_ext --inplace时,终端显示Visual Studio被自动调用,输出内容片段如下:
...>python cpp_funcs_setup.py build_ext --inplace C:\...\Python\Python312\Lib\site-packages\setuptools\_distutils\dist.py:261: UserWarning: Unknown distribution option: 'include_path' warnings.warn(msg) running build_ext building 'cpp_funcs_wrapper' extension "C:\...\... Visual Studio\2022\...\x64\cl.exe" /c /nologo /O2 /W3 /GL /DNDEBUG /MD -IC:\...\Python\Python312\include -IC:\...\Python\Python312\Include ... /EHsc /Tpcpp_funcs.cpp /Fobuild\temp.win-amd64-cpython-312\Release\cpp_funcs.obj ... Generating code Finished generating code copying build\lib.win-amd64-cpython-312\cpp_funcs_wrapper.cp312-win_amd64.pyd -> ...
更新3:编译优化的后续观察
现在我看到编译过程里,extra_link_args=[...]中的内容被自动忽略了,但还是没搞清楚怎么让C++函数的编译优化真正生效,进而提升封装后的整体性能。
备注:内容来源于stack exchange,提问作者codev

