You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

GPU环境下ONNX模型推理时出现RuntimeError问题求助

GPU环境下ONNX模型推理时出现RuntimeError问题求助

我在使用ONNX Runtime结合Optimum和Transformers做文本分类任务时,GPU环境下运行代码抛出了RuntimeError,但CPU环境下完全正常,想请教各位怎么解决这个问题。

我的代码

from optimum.onnxruntime import ORTModelForSequenceClassification
from transformers import AutoTokenizer
from optimum.pipelines import pipeline

model = ORTModelForSequenceClassification.from_pretrained("t", provider="CUDAExecutionProvider")
tokenizer = AutoTokenizer.from_pretrained("tooka")
pipe = pipeline("text-classification", model=model, tokenizer=tokenizer, accelerator="ort")
print(pipe("This is great"))

报错信息

RuntimeError: Error when binding input: There's no data transfer registered for copying tensors from Device:[DeviceType:1 MemoryType:0 DeviceId:0] to Device:[DeviceType:0 MemoryType:0 DeviceId:0]

环境信息

  • 可用的ONNX Runtime providers:['TensorrtExecutionProvider', 'CUDAExecutionProvider', 'CPUExecutionProvider']
  • 已经尝试过根据网上建议更换不同版本的ONNX,但问题依然存在。

可能的解决方案尝试

我之前遇到过类似的设备张量传输问题,给你几个可行的排查方向:

  1. 显式指定pipeline的GPU设备
    创建pipeline时加上device=0(对应你的GPU设备ID,多GPU的话按需调整),强制让整个流程在GPU上跑,避免自动调度时出现设备不匹配:

    pipe = pipeline("text-classification", model=model, tokenizer=tokenizer, accelerator="ort", device=0)
    
  2. 重新导出适配GPU的ONNX模型
    如果你的模型是自行导出的,可能导出时没考虑GPU环境。可以用Optimum的导出工具重新导出,指定CUDA设备:

    from optimum.exporters.onnx import export
    export(
        model="t",
        output="./onnx_gpu_model",
        task="text-classification",
        device="cuda"
    )
    

    之后用导出后的本地路径加载ORTModelForSequenceClassification。

  3. 调整ONNX Runtime的provider配置
    新版本Optimum推荐用providers参数指定多个provider并调整顺序,确保CUDA优先:

    model = ORTModelForSequenceClassification.from_pretrained(
        "t",
        providers=["CUDAExecutionProvider", "CPUExecutionProvider"]
    )
    

    这样可以避免多provider共存时的冲突问题。

  4. 确认版本兼容性
    虽然你换过ONNX版本,但要注意CUDA、cuDNN和ONNX Runtime GPU版本必须匹配。比如ONNX Runtime 1.15+需要CUDA 11.6及以上,建议对照官方文档确认版本组合,重新安装适配的版本。

  5. 手动测试模型GPU运行
    先跳过pipeline,直接用模型测试GPU是否能正常处理:

    inputs = tokenizer("This is great", return_tensors="pt").to("cuda")
    outputs = model(**inputs)
    print(outputs)
    

    如果这个能正常运行,说明问题出在pipeline的加速器配置上,可以进一步排查accelerator="ort"的参数设置。

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 07:18:07