You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Cython中正确使用多线程?嵌套循环代码未提速排查

问题描述

我有一段双层嵌套循环的Cython代码,想通过多线程提升速度,但修改后没看到性能提升,且只有单个线程占用率在12-14%左右。以下是相关代码和配置:

原始单线程代码

import cython
import numpy as np
from cython.parallel import prange

@cython.boundscheck(False)
@cython.wraparound(False)

cpdef int test_speed(unsigned char[:, :] Rint, unsigned char[:, :] Gint, unsigned char[:, :] Bint, unsigned char[:, :] table):
    cdef unsigned char[:, :] R_ch_new=np.zeros((1024, 1024), dtype="uint8")
    cdef unsigned char[:, :] B_ch_new=np.zeros((1024, 1024), dtype="uint8")
    cdef unsigned char[:, :] G_ch_new=np.zeros((1024, 1024), dtype="uint8")
    cdef int i, j
    for i in range(1024):
           for j in range(1024):
               # r = np.round(Rint[i,j])
               R_ch_new[i,j]=table[0, Rint[i,j]]
               # s = np.round(Gint[i,j])
               G_ch_new[i,j]=table[0, Gint[i,j]]
               # t = np.round(Bint[i,j])
               B_ch_new[i,j]=table[0, Bint[i,j]]
    return 1

多线程代码

import cython
import numpy as np
from cython.parallel import prange

@cython.boundscheck(False)
@cython.wraparound(False)
cpdef int test_speed(unsigned char[:, :] Rint, unsigned char[:, :] Gint, unsigned char[:, :] Bint, unsigned char[:, :] table):
    cdef unsigned char[:, :] R_ch_new=np.zeros((1024, 1024), dtype='uint8')
    cdef unsigned char[:, :] B_ch_new=np.zeros((1024, 1024), dtype='uint8')
    cdef unsigned char[:, :] G_ch_new=np.zeros((1024, 1024), dtype='uint8')

    cdef int i, j

    for i in prange(1024, nogil=True):
           for j in range(1024):
               # r = np.round(Rint[i,j])
               R_ch_new[i,j]=table[0, Rint[i,j]]
               # s = np.round(Gint[i,j])
               G_ch_new[i,j]=table[0, Gint[i,j]]
               # t = np.round(Bint[i,j])
               B_ch_new[i,j]=table[0, Bint[i,j]]
    return 1

测速代码

import time

import numpy as np
from testspeed1 import test_speed
table=np.zeros((1, 256), dtype='uint8')
Rint=np.zeros((1024, 1024), dtype='uint8')
Gint=np.zeros((1024, 1024), dtype='uint8')
Bint=np.zeros((1024, 1024), dtype='uint8')

for i in range(1024):
    for j in range(1024):
        Rint[i, j] = np.random.randint(0, 255)
        Gint[i, j] = np.random.randint(0, 255)
        Bint[i, j] = np.random.randint(0, 255)

for i in range(256):
    table[0, i] = np.random.randint(0, 255)

start = time.time()
result = test_speed(Rint, Gint, Bint, table)
end = time.time()
print('duration : ', str(end - start))

setup.py配置

from distutils.core import setup

import numpy
from Cython.Build import cythonize
from distutils.extension import Extension
from Cython.Distutils import build_ext

ext_modules = [
    Extension("testspeed1",
              ["testspeed1.pyx"],
              extra_compile_args=['-fopenmp'],
              include_dirs=[numpy.get_include()]
              )
]

setup(
    name="testspeed1",
    ext_modules=cythonize(ext_modules),
    include_dirs=[numpy.get_include()]
)

编译命令:python setup.py build_ext --inplace


问题分析与解决方案

1. OpenMP链接参数缺失

你的setup.py只添加了编译参数-fopenmp,但缺少链接参数。编译时必须告诉编译器链接OpenMP库,否则多线程代码不会真正并行执行,只会以单线程运行。

修改setup.py,在Extension中新增extra_link_args=['-fopenmp']:

ext_modules = [
    Extension("testspeed1",
              ["testspeed1.pyx"],
              extra_compile_args=['-fopenmp'],
              extra_link_args=['-fopenmp'],  # 新增该行
              include_dirs=[numpy.get_include()]
              )
]

2. 未显式设置线程数

默认情况下,OpenMP可能不会自动使用全部CPU核心。可以通过环境变量显式指定线程数,确保多核心被利用。

在测速代码开头添加:

import os
os.environ['OMP_NUM_THREADS'] = str(os.cpu_count())  # 使用全部CPU核心

3. 内存访问优化(可选)

当前代码对table的索引是二维形式table[0, Rint[i,j]],可以将table改为一维数组,减少索引层级,提升缓存命中率:

  • Cython函数参数改为unsigned char[:] table
  • 内部赋值语句改为R_ch_new[i,j] = table[Rint[i,j]](G、B通道同理)
  • 测速代码中table改为一维:table=np.zeros(256, dtype='uint8')

4. 验证并行效果

修改后重新编译运行,观察CPU占用率是否提升至接近100%(多核心场景),同时对比单线程与多线程的运行时间,确认性能提升。


内容的提问来源于stack exchange,提问作者sinamcr7

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 09:36:10