如何在Cython中正确使用多线程?嵌套循环代码未提速排查
问题描述
我有一段双层嵌套循环的Cython代码,想通过多线程提升速度,但修改后没看到性能提升,且只有单个线程占用率在12-14%左右。以下是相关代码和配置:
原始单线程代码
import cython import numpy as np from cython.parallel import prange @cython.boundscheck(False) @cython.wraparound(False) cpdef int test_speed(unsigned char[:, :] Rint, unsigned char[:, :] Gint, unsigned char[:, :] Bint, unsigned char[:, :] table): cdef unsigned char[:, :] R_ch_new=np.zeros((1024, 1024), dtype="uint8") cdef unsigned char[:, :] B_ch_new=np.zeros((1024, 1024), dtype="uint8") cdef unsigned char[:, :] G_ch_new=np.zeros((1024, 1024), dtype="uint8") cdef int i, j for i in range(1024): for j in range(1024): # r = np.round(Rint[i,j]) R_ch_new[i,j]=table[0, Rint[i,j]] # s = np.round(Gint[i,j]) G_ch_new[i,j]=table[0, Gint[i,j]] # t = np.round(Bint[i,j]) B_ch_new[i,j]=table[0, Bint[i,j]] return 1
多线程代码
import cython import numpy as np from cython.parallel import prange @cython.boundscheck(False) @cython.wraparound(False) cpdef int test_speed(unsigned char[:, :] Rint, unsigned char[:, :] Gint, unsigned char[:, :] Bint, unsigned char[:, :] table): cdef unsigned char[:, :] R_ch_new=np.zeros((1024, 1024), dtype='uint8') cdef unsigned char[:, :] B_ch_new=np.zeros((1024, 1024), dtype='uint8') cdef unsigned char[:, :] G_ch_new=np.zeros((1024, 1024), dtype='uint8') cdef int i, j for i in prange(1024, nogil=True): for j in range(1024): # r = np.round(Rint[i,j]) R_ch_new[i,j]=table[0, Rint[i,j]] # s = np.round(Gint[i,j]) G_ch_new[i,j]=table[0, Gint[i,j]] # t = np.round(Bint[i,j]) B_ch_new[i,j]=table[0, Bint[i,j]] return 1
测速代码
import time import numpy as np from testspeed1 import test_speed table=np.zeros((1, 256), dtype='uint8') Rint=np.zeros((1024, 1024), dtype='uint8') Gint=np.zeros((1024, 1024), dtype='uint8') Bint=np.zeros((1024, 1024), dtype='uint8') for i in range(1024): for j in range(1024): Rint[i, j] = np.random.randint(0, 255) Gint[i, j] = np.random.randint(0, 255) Bint[i, j] = np.random.randint(0, 255) for i in range(256): table[0, i] = np.random.randint(0, 255) start = time.time() result = test_speed(Rint, Gint, Bint, table) end = time.time() print('duration : ', str(end - start))
setup.py配置
from distutils.core import setup import numpy from Cython.Build import cythonize from distutils.extension import Extension from Cython.Distutils import build_ext ext_modules = [ Extension("testspeed1", ["testspeed1.pyx"], extra_compile_args=['-fopenmp'], include_dirs=[numpy.get_include()] ) ] setup( name="testspeed1", ext_modules=cythonize(ext_modules), include_dirs=[numpy.get_include()] )
编译命令:python setup.py build_ext --inplace
问题分析与解决方案
1. OpenMP链接参数缺失
你的setup.py只添加了编译参数-fopenmp,但缺少链接参数。编译时必须告诉编译器链接OpenMP库,否则多线程代码不会真正并行执行,只会以单线程运行。
修改setup.py,在Extension中新增extra_link_args=['-fopenmp']:
ext_modules = [ Extension("testspeed1", ["testspeed1.pyx"], extra_compile_args=['-fopenmp'], extra_link_args=['-fopenmp'], # 新增该行 include_dirs=[numpy.get_include()] ) ]
2. 未显式设置线程数
默认情况下,OpenMP可能不会自动使用全部CPU核心。可以通过环境变量显式指定线程数,确保多核心被利用。
在测速代码开头添加:
import os os.environ['OMP_NUM_THREADS'] = str(os.cpu_count()) # 使用全部CPU核心
3. 内存访问优化(可选)
当前代码对table的索引是二维形式table[0, Rint[i,j]],可以将table改为一维数组,减少索引层级,提升缓存命中率:
- Cython函数参数改为
unsigned char[:] table - 内部赋值语句改为
R_ch_new[i,j] = table[Rint[i,j]](G、B通道同理) - 测速代码中
table改为一维:table=np.zeros(256, dtype='uint8')
4. 验证并行效果
修改后重新编译运行,观察CPU占用率是否提升至接近100%(多核心场景),同时对比单线程与多线程的运行时间,确认性能提升。
内容的提问来源于stack exchange,提问作者sinamcr7
相关产品推荐
相关产品推荐

