You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用CuPy(Python中CUDA)时出现内存泄漏与内核停滞问题

问题:CuPy原始CUDA内核导致CPU内存飙升且内核卡住

疑问点

  1. 为何出现CPU侧内存泄漏?
  2. 为何CUDA内核始终无法执行完成?

问题复现代码

Python-CuPy版本(无法正常运行)

import numpy as np
import cupy as cp


# custom raw kernel
custom_kernel = cp.RawKernel(r'''
extern "C" __global__
void custom_kernel(double* large_array)
{
    int y = blockIdx.y * blockDim.y + threadIdx.y;
    int x = blockIdx.x * blockDim.x + threadIdx.x;
    int frame = blockIdx.z * blockDim.z + threadIdx.z;
}
''', 'custom_kernel')


# launch kernel
large_array_gpu = cp.zeros((101*101*9*9*301), dtype=cp.float64) # around 2 GB
block_dim_2 = (32, 32, 1)
bx2 = int((101 * 101 + block_dim_2[0] - 1) / block_dim_2[0])
by2 = int((9 * 9 + block_dim_2[1] - 1) / block_dim_2[1])
bz2 = int((301 + block_dim_2[2] -1 ) / block_dim_2[2])
grid_dim_2 = (bx2, by2, bz2)

custom_kernel(grid_dim_2, block_dim_2, large_array_gpu) # gets stuck at this statement, and RAM usage keeps increasing

large_array_cpu = cp.asnumpy(large_array_gpu)

print('done')

C++版本(正常运行)

#include <stdio.h>

// gpu
#include <cuda_runtime.h>
#include "device_launch_parameters.h"

__global__ void test_kernel(double* large_array)
{
    int y = blockIdx.y * blockDim.y + threadIdx.y;
    int x = blockIdx.x * blockDim.x + threadIdx.x;
    int frame = blockIdx.z * blockDim.z + threadIdx.z;

    if (y < (9 * 9) && x < (101 * 101) && frame < 301)
    {
        int resultIdx = (frame * (101 * 101) * (9 * 9)) + (y * (101 * 101) + x);
        large_array[resultIdx] = 1.1;
    }
}

int main()
{
    printf("start...");

    cudaError_t cudaStatus;

    // device
    double* dev_largeArray = 0;

    // Memory allocations   
    cudaStatus = cudaMalloc((void**)&dev_largeArray, 101 * 101 * 9 * 9 * 301 * sizeof(double));
    cudaMemset(dev_largeArray, 0, 101 * 101 * 9 * 9 * 301 * sizeof(double)); // initialize the result with zeros

    dim3 blockSize(32, 32, 1);
    int bx2 = ((101 * 101) + blockSize.x - 1) / blockSize.x;
    int by2 = ((9 * 9) + blockSize.y - 1) / blockSize.y;
    int bz2 = (301 + blockSize.z - 1) / blockSize.z;
    dim3 gridSize = dim3(bx2, by2, bz2);
    test_kernel <<<gridSize, blockSize>>> (dev_largeArray);

    // Check for any errors launching the kernel
    cudaStatus = cudaGetLastError();
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "Kernel launch failed: %s\n", cudaGetErrorString(cudaStatus));
    }

    // cudaDeviceSynchronize waits for the kernel to finish, and returns
    // any errors encountered during the launch.
    cudaStatus = cudaDeviceSynchronize();
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "cudaDeviceSynchronize returned error code %d after launching Kernel!\n", cudaStatus);
    }

    // Copy the results back to the host
    double* h_largeArray = new double[101 * 101 * 9 * 9 * 301];
    cudaStatus = cudaMemcpy(h_largeArray, dev_largeArray, 101 * 101 * 9 * 9 * 301 * sizeof(double), cudaMemcpyDeviceToHost);
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "cudaMemcpy failed!");
    }

    delete[] h_largeArray;

    cudaFree(dev_largeArray);
    return 0;
}

解答

1. CPU内存飙升的原因

你的Python代码中,CuPy的RawKernel在处理大规模线程块总数时触发了主机端内存泄漏。具体来说,你设置的grid维度为(320,4,301),总线程块数达到385280——虽然该数值在CUDA官方限制范围内,但CuPy的RawKernel调用逻辑对这种规模的线程块处理存在隐性缺陷,导致主机端持续分配内存却无法释放。

此外,Python代码中缺少C++版本那样的内核启动错误检查和同步逻辑,CuPy在启动失败后可能进入无限重试或错误循环,进一步加剧内存占用。

2. CUDA内核无法完成的原因

核心问题是线程有效性未校验+CuPy异步执行的隐性错误:

  • 你的Python内核没有像C++版本那样添加线程边界检查,大量无效线程空转,占用GPU资源但无法完成有效任务;
  • CuPy的RawKernel默认异步执行,且你未添加同步和错误检查,一旦内核启动出现隐性问题(比如参数传递的底层逻辑错误),主机端会一直等待GPU信号,表现为程序卡住。

修复方案

  1. 添加线程边界检查:在内核中加入与C++版本一致的有效性判断,避免无效线程空转:
custom_kernel = cp.RawKernel(r'''
extern "C" __global__
void custom_kernel(double* large_array)
{
    int y = blockIdx.y * blockDim.y + threadIdx.y;
    int x = blockIdx.x * blockDim.x + threadIdx.x;
    int frame = blockIdx.z * blockDim.z + threadIdx.z;

    if (y < 81 && x < 10201 && frame < 301) {
        int resultIdx = (frame * 10201 * 81) + (y * 10201 + x);
        // 可添加业务逻辑,如large_array[resultIdx] = 1.1;
    }
}
''', 'custom_kernel')
  1. 优化线程块维度:调整block的z维度,减少总线程块数,降低CuPy处理压力。例如将block_dim_2改为(32,32,8),此时bz2=(301+8-1)/8=39,总线程块数降至49920:
block_dim_2 = (32, 32, 8)
bx2 = int((101 * 101 + block_dim_2[0] - 1) / block_dim_2[0])
by2 = int((9 * 9 + block_dim_2[1] - 1) / block_dim_2[1])
bz2 = int((301 + block_dim_2[2] -1 ) / block_dim_2[2])
grid_dim_2 = (bx2, by2, bz2)
  1. 添加同步与错误检查:强制同步GPU并捕获错误,及时定位问题:
custom_kernel(grid_dim_2, block_dim_2, large_array_gpu)
cp.cuda.Device().synchronize()  # 等待内核执行完成
err = cp.cuda.get_last_error()
if err:
    print(f"CUDA error: {err}")
  1. 版本兼容检查:更新CuPy到最新稳定版,确保与你的CUDA驱动版本兼容,避免底层bug。

内容的提问来源于stack exchange,提问作者skm

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 01:02:03