You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

执行速度对比:Cython与ctypes性能差异原因探究

为什么Cython调用C比ctypes快这么多?

我正在学习Python与C的多种交互方式,实现了一个计算0到输入值间整数和的函数,分别用Python、Cython、ctypes调用自行编写的C代码实现。以输入值5000、运行1000次为测试条件,得到以下结果:

  • Python:约0.2秒
  • Ctypes:约0.01秒(快约20倍)
  • Cython:约0.0013秒(快约154倍)

Cython的执行速度远快于ctypes调用C的方式,现附上所有实现细节,请教为何ctypes方案比Cython慢这么多?


Python实现

example_py.py:

def sumTo(x):
    y = 0
    for i in range(x):
        y += i
    return y

Cython实现

example_cy.pyx:

cpdef int sumTo(int x):
    cdef int y = 0
    cdef int i
    for i in range(x):
        y += i
    return y

编译脚本setup.py:

from distutils.core import setup
from Cython.Build import cythonize

setup(ext_modules = cythonize('example_cy.pyx'))

编译命令:python setup.py build_ext --inplace

C语言实现

example_C.h:

#ifndef EXAMPLE_C_
#define EXAMPLE_C_

int sumTo(int x);

#endif

example_C.c:

#include <stdio.h>
#include "example_C.h"

int sumTo(int x) {
    int y = 0;

    for(int i = 0; i < x; i++) {
        y += i;
    }

    return y;
}

编译命令:gcc -shared -o libcalci.so -fPIC example_C.c

测试脚本

ctypes测试脚本

import example_py
from ctypes import *
import time

numRuns = 1000
x = 5000

# 测试Python脚本
tic = time.perf_counter()
for i in range(numRuns):
    example_py.sumTo(x)
py_runtime = time.perf_counter() - tic

# 测试C脚本
libCalc = CDLL("./libcalci.so")
tic = time.perf_counter()
for i in range(numRuns):
    libCalc.sumTo(x)
c_runtime = time.perf_counter() - tic

# 打印结果
print(py_runtime, c_runtime)
print('Ctypes is {}x faster'.format(py_runtime/c_runtime))

Cython测试脚本

import time
import example_cy
import example_py

numRuns = 1000
x = 5000

# 测试Python脚本
tic = time.perf_counter()
for i in range(numRuns):
    example_py.sumTo(x)
py_runtime = time.perf_counter() - tic

# 测试C脚本
tic = time.perf_counter()
for i in range(numRuns):
    example_cy.sumTo(x)
c_runtime = time.perf_counter() - tic

# 打印结果
print(py_runtime, c_runtime)
print('Cython is {}x faster'.format(py_runtime/c_runtime))

核心原因分析

  1. 跨语言调用开销差异

    • ctypes每次调用C函数都要完成Python到C的边界切换:包括参数类型校验、Python对象到C原生类型的转换、栈帧切换等操作,这些开销在1000次调用中会不断累积。
    • Cython编译生成的是Python扩展模块,cpdef函数的类型在编译期就已确定,调用时直接通过Python的C API执行,几乎没有额外的边界开销,接近原生C函数的调用效率。
  2. 编译优化程度不同

    • 你编译C库时使用的是默认选项,没有开启编译器优化(如-O2/-O3),生成的机器码优化程度低。而Cython在编译时通常会默认启用优化,编译器可以对整个函数做循环展开、常量传播等深度优化,进一步提升执行速度。
  3. 返回值处理开销

    • ctypes调用结束后,需要把C的返回值转换为Python对象,这又是一次额外的类型转换开销。
    • Cython的cpdef函数返回int类型时,能直接通过Python C API高效返回,避免了多余的转换步骤。

验证建议

  • 给C库添加优化选项重新编译:gcc -O3 -shared -o libcalci.so -fPIC example_C.c,重新测试后会发现ctypes的性能有所提升,和Cython的差距会缩小,但单次调用的边界开销依然存在。

内容的提问来源于stack exchange,提问作者rakmo97

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 20:03:16