You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

向CUDA核函数传递不同大小派生类对象数组的问题

CUDA核函数中无法正确使用多态对象指针数组

问题描述

我有一个指向抽象类A的指针数组,这些指针指向派生自A的类B、C的对象(两类对象大小不同),所有类都实现了返回自身size_t类型大小的虚函数size()。

我的实现思路:由于对象大小不同,逐个将对象拷贝到设备内存,再将这些对象的设备指针组成数组拷贝到设备,最后将该数组的设备指针传递给核函数,期望在核函数中使用这些派生类对象。

遇到的问题:使用Nsight调试器运行程序时,程序卡在printf("%d\n", objects[i]->size());行,推测objects[0]不是有效指针。

补充信息:运行环境为算力8.6的GPU,但编译时使用的是算力5.2的编译选项。

MCVE代码

#include "cuda_runtime.h"
#include "device_launch_parameters.h"

#include <stdio.h>
#include <vector>

class A {
public:
    float a;
    A() {}
    __host__ __device__ virtual size_t size() = 0;
};

class B : public A {
public:
    float b;
    B() {}
    __host__ __device__ virtual size_t size() override { return sizeof(*this); }
};

class C : public A {
public:
    float c, d;
    C() {}
    __host__ __device__ virtual size_t size() override { return sizeof(*this); }
};

__global__ void testKernel(A** objects, int numObjects) {
    for (int i = 0; i < numObjects; i++) {
        printf("%d\n", objects[i]->size());
    }
}

int main()
{
    std::vector<A*> host_pointers;
    host_pointers.push_back(new B());
    host_pointers.push_back(new C());

    cudaError_t cudaStatus;

    std::vector<A*> device_pointers;
    for (auto obj : host_pointers) {
        A* device_pointer;

        cudaStatus = cudaMalloc((void**)&device_pointer, obj->size());
        if (cudaStatus != cudaSuccess) {
            fprintf(stderr, "cudaMalloc failed for size %d\n", obj->size());
            exit(-1);
        }

        cudaStatus = cudaMemcpy(device_pointer, obj, obj->size(), cudaMemcpyHostToDevice);
        if (cudaStatus != cudaSuccess) {
            fprintf(stderr, "cudaMemcpy failed for size %d\n", obj->size());
            exit(-1);
        }
        device_pointers.push_back(device_pointer);
    }
    ///By this point, both objects should have been copied over
    ///to device memory, and I should have valid pointers to them
    A** array_of_device_pointers;

    cudaStatus = cudaMalloc((void**)&array_of_device_pointers, device_pointers.size() * sizeof(A*));
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "cudaMalloc failed\n");
        exit(-1);
    }

    cudaStatus = cudaMemcpy(array_of_device_pointers, device_pointers.data(), device_pointers.size() * sizeof(A*), cudaMemcpyHostToDevice);

    testKernel<<<1, 1>>>(array_of_device_pointers, device_pointers.size());

    cudaStatus = cudaGetLastError();
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "kernel failed, reason: %s\n", cudaGetErrorString(cudaStatus));
    }
    
    cudaStatus = cudaDeviceSynchronize();
    if (cudaStatus != cudaSuccess) {
        fprintf(stderr, "cudaSynchronize failed\n");
        exit(-1);
    }
}

问题原因与解决方法

核心原因

  1. 虚函数表指针失效:主机端创建的C++对象,其虚函数表存放在主机内存中。直接将对象拷贝到设备后,对象内的虚函数表指针仍然指向主机端地址,设备端无法访问主机内存,调用虚函数时会触发内存访问错误。
  2. 编译算力不匹配:运行环境是算力8.6的GPU,但编译时使用算力5.2的选项,可能导致设备端虚函数支持等特性不兼容,加剧问题。

解决步骤

  1. 使用统一内存或设备端构造对象
    避免直接拷贝主机对象到设备,改用CUDA统一内存(Unified Memory)管理对象,让CUDA自动处理内存映射和虚函数表的设备端适配。示例修改如下:

    std::vector<A*> host_pointers;
    B* b_ptr;
    // 分配统一内存
    cudaMallocManaged(&b_ptr, sizeof(B));
    // 在统一内存中原地构造对象
    new(b_ptr) B();
    host_pointers.push_back(b_ptr);
    
    C* c_ptr;
    cudaMallocManaged(&c_ptr, sizeof(C));
    new(c_ptr) C();
    host_pointers.push_back(c_ptr);
    

    之后无需手动拷贝对象到设备,直接将指针数组拷贝到设备即可,统一内存会自动在主机和设备间同步数据。

  2. 匹配编译算力
    编译时添加对应算力的编译选项,比如针对算力8.6的GPU,添加-arch=sm_86,确保生成的代码与运行环境兼容。

  3. 验证指针有效性
    在核函数中先打印设备指针地址,确认指针是否有效:

    __global__ void testKernel(A** objects, int numObjects) {
        for (int i = 0; i < numObjects; i++) {
            printf("Device pointer %d: %p\n", i, objects[i]);
            printf("Size: %zu\n", objects[i]->size());
        }
    }
    

内容的提问来源于stack exchange,提问作者Hessian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 15:45:37