执行速度对比:Cython与ctypes性能差异原因探究
为什么Cython调用C比ctypes快这么多?
我正在学习Python与C的多种交互方式,实现了一个计算0到输入值间整数和的函数,分别用Python、Cython、ctypes调用自行编写的C代码实现。以输入值5000、运行1000次为测试条件,得到以下结果:
- Python:约0.2秒
- Ctypes:约0.01秒(快约20倍)
- Cython:约0.0013秒(快约154倍)
Cython的执行速度远快于ctypes调用C的方式,现附上所有实现细节,请教为何ctypes方案比Cython慢这么多?
Python实现
example_py.py:
def sumTo(x): y = 0 for i in range(x): y += i return y
Cython实现
example_cy.pyx:
cpdef int sumTo(int x): cdef int y = 0 cdef int i for i in range(x): y += i return y
编译脚本setup.py:
from distutils.core import setup from Cython.Build import cythonize setup(ext_modules = cythonize('example_cy.pyx'))
编译命令:python setup.py build_ext --inplace
C语言实现
example_C.h:
#ifndef EXAMPLE_C_ #define EXAMPLE_C_ int sumTo(int x); #endif
example_C.c:
#include <stdio.h> #include "example_C.h" int sumTo(int x) { int y = 0; for(int i = 0; i < x; i++) { y += i; } return y; }
编译命令:gcc -shared -o libcalci.so -fPIC example_C.c
测试脚本
ctypes测试脚本
import example_py from ctypes import * import time numRuns = 1000 x = 5000 # 测试Python脚本 tic = time.perf_counter() for i in range(numRuns): example_py.sumTo(x) py_runtime = time.perf_counter() - tic # 测试C脚本 libCalc = CDLL("./libcalci.so") tic = time.perf_counter() for i in range(numRuns): libCalc.sumTo(x) c_runtime = time.perf_counter() - tic # 打印结果 print(py_runtime, c_runtime) print('Ctypes is {}x faster'.format(py_runtime/c_runtime))
Cython测试脚本
import time import example_cy import example_py numRuns = 1000 x = 5000 # 测试Python脚本 tic = time.perf_counter() for i in range(numRuns): example_py.sumTo(x) py_runtime = time.perf_counter() - tic # 测试C脚本 tic = time.perf_counter() for i in range(numRuns): example_cy.sumTo(x) c_runtime = time.perf_counter() - tic # 打印结果 print(py_runtime, c_runtime) print('Cython is {}x faster'.format(py_runtime/c_runtime))
核心原因分析
跨语言调用开销差异
- ctypes每次调用C函数都要完成Python到C的边界切换:包括参数类型校验、Python对象到C原生类型的转换、栈帧切换等操作,这些开销在1000次调用中会不断累积。
- Cython编译生成的是Python扩展模块,
cpdef函数的类型在编译期就已确定,调用时直接通过Python的C API执行,几乎没有额外的边界开销,接近原生C函数的调用效率。
编译优化程度不同
- 你编译C库时使用的是默认选项,没有开启编译器优化(如
-O2/-O3),生成的机器码优化程度低。而Cython在编译时通常会默认启用优化,编译器可以对整个函数做循环展开、常量传播等深度优化,进一步提升执行速度。
- 你编译C库时使用的是默认选项,没有开启编译器优化(如
返回值处理开销
- ctypes调用结束后,需要把C的返回值转换为Python对象,这又是一次额外的类型转换开销。
- Cython的
cpdef函数返回int类型时,能直接通过Python C API高效返回,避免了多余的转换步骤。
验证建议
- 给C库添加优化选项重新编译:
gcc -O3 -shared -o libcalci.so -fPIC example_C.c,重新测试后会发现ctypes的性能有所提升,和Cython的差距会缩小,但单次调用的边界开销依然存在。
内容的提问来源于stack exchange,提问作者rakmo97
相关产品推荐
相关产品推荐

