CuPy GPU端3D FFT卷积连续多次顺序调用后运行速度下降
问题背景
您好,我首次在技术社区提问,将尽量遵守社区发帖规范,若发帖内容有需调整之处欢迎指出。
我当前正在开发可快速从二值3D图像中提取孔径分布的函数,实现逻辑参考ImageJ的local thickness插件,通过计算图像局部厚度得到结果。由于该函数需要在模拟退火流程中被调用约20万次,理想状态下单次运行耗时需控制在1秒以内。函数运行采用CPU+GPU异构架构,测试硬件环境为:
- CPU:12代Intel Core i7-12700KF(20核,16GB内存)
- GPU:NVIDIA RTX 3050(8GB显存)
目前函数可正常输出结果,但疑似后端存在异常机制导致运行速度被拖慢,初步怀疑问题与线程调度、GPU-CPU数据交互开销、GPU温控降频的“冷却周期”相关。
现有函数执行流程
函数整体分为三个执行步骤:
- 欧式距离变换(Euclidean Distance Transform):在CPU端调用edt包并行实现,针对250³尺寸的二值图像,该步骤耗时约0.25秒;
- 3D骨架化(3d Skeletonization):在CPU端调用
skimage.morphology.skeletonize_3d实现,通过dask将图像分块处理,分块逻辑基于porespy.filters.chunked_func封装;将得到的骨架与距离变换结果相乘,即可得到取值为距最近背景体素最小距离的骨架图,该步骤耗时约0.45-0.5秒; - 骨架体素膨胀操作:对骨架上的每个体素,使用半径等于该体素取值的球形结构元进行膨胀,循环从最大结构元尺寸开始递减执行,保证大球生成的区域不会被后续小球覆盖;膨胀操作在GPU端通过
cupyx.scipy.signal.signaltools.convolve的FFT卷积实现,单次卷积耗时约0.005秒。
性能异常最小复现代码
该性能异常问题无需运行完整函数即可复现,核心触发条件为连续顺序执行多次FFT卷积,最小复现代码如下:
import skimage import time import cupy as cp from cupyx.scipy.signal.signaltools import convolve # Generate a binary image im = cp.random.random((250,250,250)) > 0.4 # Generate spherical structuring kernels for input to convolution structuring_kernels = {} for r in range(1,21): structuring_kernels.update({r: cp.array(skimage.morphology.ball(r))}) # run dilation process in loop for i in range(10): s = time.perf_counter() for j in range(20,0,-1): convolve(im, structuring_kernels[j], mode='same', method='fft') e = time.perf_counter() # time.sleep(2) print(e-s)
已观测到的测试现象
- 直接运行上述代码时,前2-3次循环后,单次完整膨胀循环在测试机上耗时约1.8秒;
- 如果取消
time.sleep(2)行的注释(即每轮外层循环之间暂停2秒),单次外层循环耗时仅为0.05秒; - 耗时是经过数轮循环后才逐步上升至1.8秒并保持稳定,监控GPU状态时可见3D负载快速飙升至100%并维持在高位;
- 补充测试:理论上完整局部厚度函数单次计算耗时应约为0.8秒,但在循环中处理不同图像时,单次耗时上升至约1.5秒;如果在每次函数调用后添加
time.sleep(1)的等待逻辑,函数单次耗时即可回到约0.8秒的预期水平。
待解答疑问
- 如果性能瓶颈来自GPU硬件算力上限,为何前几轮循环的运行速度更快?
- 该现象是否由显存泄漏导致?
- 该问题的根本成因是什么?
- 是否可以通过CuPy的后端控制接口配置来规避该性能下降问题?
完整局部厚度计算函数代码
import porespy as ps from skimage.morphology import skeletonize_3d import time import numpy as np import cupy as cp from edt import edt from cupyx.scipy.signal.signaltools import convolve def local_thickness_cp(im, masks=None, method='fft'): """ Parameters ---------- im: 3D voxelized image for which the local thickness map is desired masks: (optional) A dictionary of the structuring elements to be used method: 'fft' or 'direct' Returns ------- The local thickness map """ s = time.perf_counter() # Calculate the euclidean distance transform using edt package dt = cp.array(edt(im, parallel=15)) e = time.perf_counter() # print(f'EDT took {e - s}') s = time.perf_counter() # Calculate the skeleton of the image and multiply by dt skel = cp.array(ps.filters.chunked_func(skeletonize_3d, overlap=17, divs=[2, 3, 3], cores=20, image=im).astype(bool)) * dt e = time.perf_counter() # print(f'skeletonization took {e - s} seconds') r_max = int(cp.max(skel)) s = time.perf_counter() if not masks: masks = {} for r in range(int(r_max), 0, -1): masks.update({r: cp.array(ps.tools.ps_ball(r))}) e = time.perf_counter() # print(f'mask creation took {e - s} seconds') # Initialize the local thickness image final = cp.zeros(cp.shape(skel)) time_in_loop = 0 s = time.perf_counter() for r in range(r_max, 0, -1): # Get a mask of where the skeleton has values between r-1 and r skel_selected = ((skel > r - 1) * (skel <= r)).astype(int) # Perform dilation on the mask using fft convolve method, and multiply by radius of pore size dilation = (convolve(skel_selected, masks[r], mode='same', method=method) > 0.1) * r # Add dilation to local thickness image, where it is still zero (ie don't overwrite previous inserted values) final = final + (final == 0) * dilation e = time.perf_counter() # print(f'Dilation loop took {e - s} seconds') return final
内容的提问来源于stack exchange,提问作者Pascal R
相关产品推荐
相关产品推荐

