You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Cython中快速处理int数组的性能疑惑及最优方案问询

问题:Cython中Python列表为何比NumPy/memoryview数组更快?

我在项目中需要快速处理int数组,重点是快速创建和遍历。查资料得知NumPy数组可作为Cython typed memoryviews使用,于是写了以下测试代码:

%%cython
import numpy as np

cpdef test_1(int[:] lst):
    cdef int[100] ret
    for ii in range(lst.shape[0]):
        ret[ii] = lst[ii]
    return ret

cpdef test_2(int[:] lst):
    cdef int[:] ret = np.empty(100, dtype=np.dtype("i"))
    for ii in range(lst.shape[0]):
        ret[ii] = lst[ii]
    return ret

cpdef test_3(list lst):
    cdef list ret = [0]*100
    for ii in range(len(lst)):
        ret[ii] = lst[ii]
    return ret

预期是test_1(typed数组)最快,test_2(NumPy数组)次之,test_3(Python列表)最慢,但测试结果却相反:

import numpy as np
lst_1 = [1, 2, 3, 4, 5]
lst_2 = np.array(lst_1, dtype=np.intc)
%timeit test_1(lst_2)
%timeit test_2(lst_2)
%timeit test_3(lst_1)

1.3 µs ± 60.3 ns per loop (mean ± std. dev. of 7 runs, 1000000 loops each)
1.43 µs ± 18.3 ns per loop (mean ± std. dev. of 7 runs, 1000000 loops each)
481 ns ± 13.1 ns per loop (mean ± std. dev. of 7 runs, 1000000 loops each)

请问:

  1. 为何Python列表反而最快?
  2. 我在创建或访问memoryviews时是否存在错误?
  3. 如何在Cython中超快地创建和访问int数组?

解答

一、Python列表更快的核心原因

你的测试用例数据量太小(仅5个元素),此时Python列表的低开销优势盖过了静态类型数组的遍历效率优势:

  • test_3里的[0]*100是Python层面的批量初始化操作,底层是C级别的内存分配和填充,速度极快;而遍历仅5次,Python对象访问的开销几乎可以忽略。
  • test_1返回的是C静态数组,Cython需要自动把它转换成Python列表返回,这个转换过程要遍历100个元素并逐个封装成Python int对象,开销远大于5次遍历的时间。
  • test_2需要调用np.empty创建NumPy数组,涉及NumPy对象初始化、内存分配等额外开销,在小数据量下占比极高。

二、代码中的潜在问题

  1. 返回值转换开销:C静态数组int[100]无法直接作为Python对象返回,Cython的自动转换过程会产生大量额外开销。
  2. 数据量不匹配:输入数组仅5个元素,但目标数组是100个元素,遍历只执行5次,初始化的开销成为性能主导,而非遍历效率。
  3. 默认边界检查:Cython对memoryview的访问默认会做边界检查,虽然可以关闭,但在小数据量下影响有限。

三、Cython中超快创建和访问int数组的方案

1. 关闭边界与负索引检查

在函数顶部添加装饰器,消除内存访问的额外检查:

%%cython
import numpy as np
cimport cython

@cython.boundscheck(False)
@cython.wraparound(False)
cpdef test_1_opt(int[:] lst):
    cdef int[100] ret
    cdef int i, n = lst.shape[0]
    for i in range(n):
        ret[i] = lst[i]
    # 手动返回NumPy数组,避免自动转换为列表的开销
    return np.array(ret, dtype=np.intc)

@cython.boundscheck(False)
@cython.wraparound(False)
cpdef test_2_opt(int[:] lst):
    cdef int[:] ret = np.empty(100, dtype=np.intc)
    cdef int i, n = lst.shape[0]
    for i in range(n):
        ret[i] = lst[i]
    return ret.base  # 返回底层NumPy数组,跳过memoryview包装开销

2. 放大数据量,凸显静态数组优势

当处理大数组(比如10000+元素)时,静态类型数组的效率会显著超过Python列表。此时遍历的性能优势会盖过初始化的开销,NumPy/memoryview的速度会是Python列表的数倍。

3. 直接操作C级内存,避免Python对象转换

如果不需要返回Python对象,而是在Cython内部持续处理数组,直接使用C静态数组或memoryview,完全跳过Python对象转换的开销,这是最快的方式:

%%cython
cimport cython

@cython.boundscheck(False)
@cython.wraparound(False)
cpdef void process_in_place(int[:] arr):
    cdef int i, n = arr.shape[0]
    for i in range(n):
        arr[i] *= 2  # 直接修改内存,无Python对象交互

4. 使用Cython的array模块替代NumPy

对于小型固定大小数组,Cython的array模块开销比NumPy更小:

%%cython
from cpython cimport array
import array

cpdef test_array():
    cdef array.array ret = array.array('i', [0]*100)
    cdef int[:] ret_view = ret
    # 直接操作view
    for i in range(5):
        ret_view[i] = i+1
    return ret

内容的提问来源于stack exchange,提问作者Lewwwer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 06:50:00