Python线程无法充分利用CPU核心,如何实现满负载运行?
问题:Python多线程反投影计算无法充分利用CPU
问题描述
编写作业计算脚本时,因反投影计算特性在笔记本上运行缓慢,尝试用threading模块将主循环拆分为8个线程(与CPU核心数一致),但htop显示每个核心仅约20%负载,计算速度未达预期。线程间无共享写入,仅读取sinogram,理论适合并行,但CPU利用率上不去。
原主线程代码:
import numpy as np import matplotlib.pyplot as plt import os import threading as thr plt.close('all') os.system("sudo renice -n -19 -p " + str(os.getpid())) sinogram = np.random.rand((120, 240)) # for illustration purposes N = np.shape(sinogram)[1] # width M = np.shape(sinogram)[0] threads = [] nthreads = 8 results = [None] * nthreads for i in range(nthreads): srange = range(int((i)*M/nthreads), int((i+1)*M/nthreads)) t = thr.Thread(target=backproject_threaded, args=(sinogram,srange,results,i)) print("started thread " + str(i) + " with srange " + str(srange)) t.start() threads.append(t) backarray = np.zeros([N,N]) for i in range(nthreads): threads[i].join() backarray += results[i] plt.imshow((backarray), cmap='gray')
backproject_threaded函数:
def backproject_threaded(sinogram, srange, results, index): N = np.shape(sinogram)[1] # width M = np.shape(sinogram)[0] # height (number of projections) backarray = np.zeros([N,N]) for s in srange: angle = np.pi*s/M # angle of axis for y in range(N): for x in range(N): # project coordinate onto rotating axis z = np.cos(angle) * (x-N/2) + np.sin(angle) * (y-N/2) # axis-coordinate # select the closest coordinate in sinogram(s, :) sn = max(0, min(int((z+N/2)), N-1)) # add this value to backarray value backarray[N-y-1, x] += sinogram[s,sn] results[index] = backarray print("thread " + str(index) + " stopped") return backarray
原因分析
Python的**全局解释器锁(GIL)**是核心问题:threading模块的线程属于同一进程,共享GIL,同一时间只有一个线程能执行Python字节码。对于CPU密集型任务(如你的三重嵌套循环),多线程无法实现真正的并行,只能交替执行,导致CPU利用率极低。
解决方案
1. 改用multiprocessing实现多进程并行
多进程会创建独立的Python解释器进程,每个进程有自己的GIL,能真正利用多核CPU并行计算。修改代码如下:
import numpy as np import matplotlib.pyplot as plt import os import multiprocessing as mp plt.close('all') os.system("sudo renice -n -19 -p " + str(os.getpid())) sinogram = np.random.rand(120, 240) # 注意:rand参数不需要额外括号 N = sinogram.shape[1] M = sinogram.shape[0] nprocesses = 8 def backproject_process(sinogram, srange): N = sinogram.shape[1] backarray = np.zeros([N,N]) for s in srange: angle = np.pi*s/M for y in range(N): for x in range(N): z = np.cos(angle) * (x-N/2) + np.sin(angle) * (y-N/2) sn = max(0, min(int((z+N/2)), N-1)) backarray[N-y-1, x] += sinogram[s,sn] return backarray # 拆分任务 tasks = [] for i in range(nprocesses): start = int(i*M/nprocesses) end = int((i+1)*M/nprocesses) tasks.append((sinogram, range(start, end))) # 启动进程池 with mp.Pool(processes=nprocesses) as pool: results = pool.starmap(backproject_process, tasks) # 合并结果 backarray = sum(results) plt.imshow(backarray, cmap='gray') plt.show()
2. 向量化优化核心计算(更关键)
你的三重嵌套循环是Python层面的,即使并行也有很大开销。用NumPy向量化操作替代循环,能大幅提升单进程性能,再结合多进程效果更佳:
def backproject_vectorized(sinogram, srange): N = sinogram.shape[1] M = sinogram.shape[0] backarray = np.zeros([N, N]) # 创建网格坐标 x = np.arange(N) - N/2 y = np.arange(N) - N/2 X, Y = np.meshgrid(x, y) for s in srange: angle = np.pi * s / M # 向量化计算z z = np.cos(angle) * X + np.sin(angle) * Y # 计算sn,向量化处理边界 sn = np.round(z + N/2).astype(int) sn = np.clip(sn, 0, N-1) # 提取sinogram对应值并累加 backarray += sinogram[s, sn] # 调整y轴方向(原代码中N-y-1对应翻转) backarray = np.flipud(backarray) return backarray
将这个向量化函数替换到多进程代码中,性能会有质的提升,CPU利用率也会接近100%。
3. 其他注意事项
- 避免在计算循环中使用
print,IO操作会拖慢计算速度。 sinogram在多进程中会被复制到每个进程,若数据量极大,可考虑用mp.Array共享内存,但小数据量下无需额外处理。
内容的提问来源于stack exchange,提问作者Simon Innerbichler
相关产品推荐
相关产品推荐

