You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CuPy GPU端3D FFT卷积连续多次顺序调用后运行速度下降

问题背景

您好,我首次在技术社区提问,将尽量遵守社区发帖规范,若发帖内容有需调整之处欢迎指出。
我当前正在开发可快速从二值3D图像中提取孔径分布的函数,实现逻辑参考ImageJ的local thickness插件,通过计算图像局部厚度得到结果。由于该函数需要在模拟退火流程中被调用约20万次,理想状态下单次运行耗时需控制在1秒以内。函数运行采用CPU+GPU异构架构,测试硬件环境为:

  • CPU:12代Intel Core i7-12700KF(20核,16GB内存)
  • GPU:NVIDIA RTX 3050(8GB显存)

目前函数可正常输出结果,但疑似后端存在异常机制导致运行速度被拖慢,初步怀疑问题与线程调度、GPU-CPU数据交互开销、GPU温控降频的“冷却周期”相关。

现有函数执行流程

函数整体分为三个执行步骤:

  1. 欧式距离变换(Euclidean Distance Transform):在CPU端调用edt包并行实现,针对250³尺寸的二值图像,该步骤耗时约0.25秒;
  2. 3D骨架化(3d Skeletonization):在CPU端调用skimage.morphology.skeletonize_3d实现,通过dask将图像分块处理,分块逻辑基于porespy.filters.chunked_func封装;将得到的骨架与距离变换结果相乘,即可得到取值为距最近背景体素最小距离的骨架图,该步骤耗时约0.45-0.5秒;
  3. 骨架体素膨胀操作:对骨架上的每个体素,使用半径等于该体素取值的球形结构元进行膨胀,循环从最大结构元尺寸开始递减执行,保证大球生成的区域不会被后续小球覆盖;膨胀操作在GPU端通过cupyx.scipy.signal.signaltools.convolve的FFT卷积实现,单次卷积耗时约0.005秒。
性能异常最小复现代码

该性能异常问题无需运行完整函数即可复现,核心触发条件为连续顺序执行多次FFT卷积,最小复现代码如下:

import skimage
import time
import cupy as cp
from cupyx.scipy.signal.signaltools import convolve


# Generate a binary image
im = cp.random.random((250,250,250)) > 0.4

# Generate spherical structuring kernels for input to convolution
structuring_kernels = {}
for r in range(1,21):
    structuring_kernels.update({r: cp.array(skimage.morphology.ball(r))})

# run dilation process in loop
for i in range(10):
    s = time.perf_counter()
    for j in range(20,0,-1):
        convolve(im, structuring_kernels[j], mode='same', method='fft')
    e = time.perf_counter()
    # time.sleep(2)
    print(e-s)
已观测到的测试现象
  • 直接运行上述代码时,前2-3次循环后,单次完整膨胀循环在测试机上耗时约1.8秒;
  • 如果取消time.sleep(2)行的注释(即每轮外层循环之间暂停2秒),单次外层循环耗时仅为0.05秒;
  • 耗时是经过数轮循环后才逐步上升至1.8秒并保持稳定,监控GPU状态时可见3D负载快速飙升至100%并维持在高位;
  • 补充测试:理论上完整局部厚度函数单次计算耗时应约为0.8秒,但在循环中处理不同图像时,单次耗时上升至约1.5秒;如果在每次函数调用后添加time.sleep(1)的等待逻辑,函数单次耗时即可回到约0.8秒的预期水平。
待解答疑问
  • 如果性能瓶颈来自GPU硬件算力上限,为何前几轮循环的运行速度更快?
  • 该现象是否由显存泄漏导致?
  • 该问题的根本成因是什么?
  • 是否可以通过CuPy的后端控制接口配置来规避该性能下降问题?
完整局部厚度计算函数代码
import porespy as ps
from skimage.morphology import skeletonize_3d
import time
import numpy as np
import cupy as cp
from edt import edt
from cupyx.scipy.signal.signaltools import convolve

def local_thickness_cp(im, masks=None, method='fft'):
    """

    Parameters
    ----------
    im: 3D voxelized image for which the local thickness map is desired
    masks: (optional) A dictionary of the structuring elements to be used
    method: 'fft' or 'direct'

    Returns
    -------
    The local thickness map
    """
    s = time.perf_counter()
    # Calculate the euclidean distance transform using edt package
    dt = cp.array(edt(im, parallel=15))
    e = time.perf_counter()
    # print(f'EDT took {e - s}')
    s = time.perf_counter()
    # Calculate the skeleton of the image and multiply by dt
    skel = cp.array(ps.filters.chunked_func(skeletonize_3d,
                                            overlap=17,
                                            divs=[2, 3, 3],
                                            cores=20,
                                            image=im).astype(bool)) * dt
    e = time.perf_counter()
    # print(f'skeletonization took {e - s} seconds')
    r_max = int(cp.max(skel))
    s = time.perf_counter()
    if not masks:
        masks = {}
        for r in range(int(r_max), 0, -1):
            masks.update({r: cp.array(ps.tools.ps_ball(r))})
    e = time.perf_counter()
    # print(f'mask creation took {e - s} seconds')
    # Initialize the local thickness image
    final = cp.zeros(cp.shape(skel))
    time_in_loop = 0
    s = time.perf_counter()
    for r in range(r_max, 0, -1):
        # Get a mask of where the skeleton has values between r-1 and r
        skel_selected = ((skel > r - 1) * (skel <= r)).astype(int)
        # Perform dilation on the mask using fft convolve method, and multiply by radius of pore size
        dilation = (convolve(skel_selected, masks[r], mode='same', method=method) > 0.1) * r
        # Add dilation to local thickness image, where it is still zero (ie don't overwrite previous inserted values)
        final = final + (final == 0) * dilation

    e = time.perf_counter()
    # print(f'Dilation loop took {e - s} seconds')
    return final

内容的提问来源于stack exchange,提问作者Pascal R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 02:57:14