You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Colab中无法安装dgl-cu库及模型训练报错求助

解决Google Colab中DGL GPU版本安装失败及训练报错问题

我在Google Colab使用A100 GPU运行DGL图模型时,训练阶段持续报错,核心问题是无法通过!pip install dgl-cu111安装DGL的CUDA版本,安装时出现如下错误:

Looking in indexes: https://pypi.org/simple, https://us-python.pkg.dev/colab-wheels/public/simple/
ERROR: Could not find a version that satisfies the requirement dgl-cu111 (from versions: none)
ERROR: No matching distribution found for dgl-cu111

训练时的报错信息:

load_done
/usr/local/lib/python3.10/dist-packages/torch/nn/modules/rnn.py:71: UserWarning: dropout option adds dropout after all but last recurrent layer, so non-zero dropout expects num_layers greater than 1, but got dropout=0.2 and num_layers=1
  warnings.warn("dropout option adds dropout after all but last "
init done
Epoch 1:   0%|          | 0/49 [00:00<?, ?it/s]
---------------------------------------------------------------------------
DGLError                                  Traceback (most recent call last)
<ipython-input-7-783797e86ab0> in <cell line: 209>()
    207 
    208 
--> 209 train(model)
    210 test_func(model, y_test, X_test)

3 frames
<ipython-input-7-783797e86ab0> in train(net)
    179                         gc.collect()
    180                         continue
--> 181                     acc, loss, _ = fwd_pass(batch_X, batch_y, train=True)
    182 
    183                     losses.append(loss.item())

<ipython-input-7-783797e86ab0> in fwd_pass(X, y, train)
    108     for item in X:
    109         x = [0, 0]
--> 110         x[0] = item[0].to(device)
    111         x[1] = item[1].to(device)
    112         out.append(model(x))

/usr/local/lib/python3.10/dist-packages/dgl/heterograph.py in to(self, device, **kwargs)
   5707 
   5708         # 1. Copy graph structure
--> 5709         ret._graph = self._graph.copy_to(utils.to_dgl_context(device))
   5710 
   5711         # 2. Copy features

/usr/local/lib/python3.10/dist-packages/dgl/heterograph_index.py in copy_to(self, ctx)
    253             The graph index on the given device context.
    254         """
--> 255         return _CAPI_DGLHeteroCopyTo(self, ctx.device_type, ctx.device_id)
    256 
    257     def pin_memory(self):

dgl/_ffi/_cython/./function.pxi in dgl._ffi._cy3.core.FunctionBase.__call__()

dgl/_ffi/_cython/./function.pxi in dgl._ffi._cy3.core.FuncCall()

dgl/_ffi/_cython/./function.pxi in dgl._ffi._cy3.core.FuncCall3()

DGLError: [01:00:58] /opt/dgl/src/runtime/c_runtime_api.cc:82: Check failed: allow_missing: Device API cuda is not enabled. Please install the cuda version of dgl.
Stack trace:
  [bt] (0) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dmlc::LogMessageFatal::~LogMessageFatal()+0x75) [0x7fae2b978e55]
  [bt] (1) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::runtime::DeviceAPIManager::GetAPI(std::string, bool)+0x1f2) [0x7fae2bcf85f2]
  [bt] (2) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::runtime::DeviceAPI::Get(DGLContext, bool)+0x1e1) [0x7fae2bcf2ba1]
  [bt] (3) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::runtime::NDArray::Empty(std::vector<long, std::allocator<long> >, DGLDataType, DGLContext)+0x13b) [0x7fae2bd15acb]
  [bt] (4) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::runtime::NDArray::CopyTo(DGLContext const&) const+0xc3) [0x7fae2bd4fe23]
  [bt] (5) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::UnitGraph::CopyTo(std::shared_ptr<dgl::BaseHeteroGraph>, DGLContext const&)+0x3ef) [0x7fae2be5d79f]
  [bt] (6) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(dgl::HeteroGraph::CopyTo(std::shared_ptr<dgl::BaseHeteroGraph>, DGLContext const&)+0xf6) [0x7fae2bd61286]
  [bt] (7) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(+0x52cbb6) [0x7fae2bd70bb6]
  [bt] (8) /usr/local/lib/python3.10/dist-packages/dgl/libdgl.so(DGLFuncCall+0x48) [0x7fae2bcf7bb8]

当前Colab环境的CUDA版本:

nvcc: NVIDIA (R) Cuda compiler driver
Copyright (c) 2005-2022 NVIDIA Corporation
Built on Wed_Sep_21_10:33:58_PDT_2022
Cuda compilation tools, release 11.8, V11.8.89
Build cuda_11.8.r11.8/compiler.31833905_0

解决方案

核心问题是安装的DGL版本和当前CUDA版本不匹配:你尝试安装dgl-cu111,但实际环境是CUDA 11.8,DGL的CUDA版本包命名严格对应CUDA主版本号,所以需要安装dgl-cu118。

执行以下步骤:

  1. 先卸载已安装的CPU版DGL(如果有的话):
!pip uninstall -y dgl
  1. 安装对应CUDA 11.8的DGL版本:
!pip install dgl-cu118 -f https://data.dgl.ai/wheels/cu118/repo.html
  1. 验证安装是否成功:
import dgl
print(dgl.backend.backend_name())
# 输出应为 'pytorch'(如果你用PyTorch),且运行dgl.device('cuda')不会报错
print(dgl.device('cuda'))

执行完以上步骤后,重新运行你的训练代码,应该就能解决CUDA未启用的报错了。


内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 03:17:14