You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用thrust::transform()从主机到设备时触发cudaErrorIllegalAddress错误

问题:Thrust::Transform主机到设备转换触发非法内存访问错误

执行thrust::transform(h_vec.cbegin(), h_vec.cend(), device_data.begin(), Cast<DEVICE_TYPE>{})语句时触发cudaErrorIllegalAddress非法内存访问错误,程序崩溃终止。但类似的thrust::copy()语句可正常完成主机到设备的数据拷贝,尝试使用thrust::host或thrust::device执行策略也未解决问题(CUDA版本12.3,实际应用中HOST_TYPE为char,调试时改为int32_t)。

程序代码

#include <thrust/copy.h>
#include <thrust/execution_policy.h>
#include <thrust/transform.h>
#include <thrust/device_vector.h>
#include <thrust/host_vector.h>
#include <iostream>

using HOST_TYPE=int32_t;
using DEVICE_TYPE=int;

template <typename T>
struct Cast {
  __host__ __device__ T operator()(HOST_TYPE i) const {
    return static_cast<T>(i);
  }
};

int main() {
    // Initialize host data
    thrust::host_vector<HOST_TYPE> const h_vec{1, 2, 3, 4, 5};
    
    // Allocate space on the device
    thrust::device_vector<DEVICE_TYPE> device_data(h_vec.size());

    // Copy data from host to device
    //thrust::copy(h_vec.cbegin(), h_vec.cend(), device_data.begin());  // this works
    thrust::transform(h_vec.cbegin(), h_vec.cend(), device_data.begin(), Cast<DEVICE_TYPE>{});
    
    // Copy back to host to check
    thrust::host_vector<DEVICE_TYPE> host_data_copy = device_data;
    for (DEVICE_TYPE val : host_data_copy) {
        std::cout << val << " ";
    }
    std::cout << std::endl;
    
    return 0;
}

错误信息

$ nvcc test.cu
$ ./a.out 
terminate called after throwing an instance of 'thrust::system::system_error'
  what():  parallel_for: failed to synchronize: cudaErrorIllegalAddress: an illegal memory access was encountered
Aborted (core dumped)

原因分析与解决方法

核心原因

Thrust的transform操作默认会选择在设备端执行,但此时输入迭代器指向主机内存(host_vector的迭代器),设备端代码无法直接访问主机内存,因此触发非法内存访问错误。而thrust::copy有专门的跨内存空间处理逻辑,会自动完成主机到设备的数据拷贝,无需额外配置。

如果之前尝试thrust::host执行策略无效,大概率是未将执行策略作为第一个参数传入transform,导致策略未生效。

解决方法

方法1:显式指定主机执行策略

将thrust::host作为第一个参数传入transform,让转换操作在主机端执行,Thrust会自动将结果拷贝到设备向量:

thrust::transform(thrust::host, h_vec.cbegin(), h_vec.cend(), device_data.begin(), Cast<DEVICE_TYPE>{});

方法2:先拷贝主机数据到设备,再执行设备端transform

先将主机数据拷贝到临时设备向量,再在设备端完成转换,避免跨内存空间访问:

thrust::device_vector<HOST_TYPE> temp_device(h_vec);
thrust::transform(temp_device.cbegin(), temp_device.cend(), device_data.begin(), Cast<DEVICE_TYPE>{});

内容的提问来源于stack exchange,提问作者Matt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 00:12:52