You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

torch.normal().to('cuda')报CUDA非法内存访问错误的解决求助

环境

Package                 Version
----------------------- ------------
absl-py 2.0.0
blessed 1.20.0
cachetools 5.3.2
captum 0.6.0
certifi 2022.12.7
charset-normalizer 3.3.1
colour 0.1.5
cycler 0.11.0
Cython 3.0.5
dlib 19.24.0
fonttools 4.38.0
future 0.18.3
google-auth 2.23.4
google-auth-oauthlib 0.4.6
gpustat 1.1.1
grpcio 1.59.2
idna 3.4
importlib-metadata 6.7.0
joblib 1.3.2
kiwisolver 1.4.5
Markdown 3.4.4
MarkupSafe 2.1.3
matplotlib 3.5.3
numpy 1.21.6
nvidia-ml-py 12.535.108
oauthlib 3.2.2
packaging 23.2
pandas 1.1.5
patsy 0.5.3
Pillow 9.5.0
pip 22.3.1
protobuf 3.20.3
psutil 5.9.6
pyasn1 0.5.0
pyasn1-modules 0.3.0
pycocotools 2.0.7
pyparsing 3.1.1
python-dateutil 2.8.2
pytz 2023.3.post1
requests 2.31.0
requests-oauthlib 1.3.1
rsa 4.9
rtpt 0.0.4
scikit-learn 1.0.2
scipy 1.7.3
seaborn 0.12.2
setproctitle 1.3.3
setuptools 65.6.3
six 1.16.0
statsmodels 0.13.5
tensorboard 2.11.2
tensorboard-data-server 0.6.1
tensorboard-logger 0.1.0
tensorboard-plugin-wit 1.8.1
Tensorboard dx 2.6.2.2
threadpoolctl 3.1.0
torch 1.6.0+cu101
torchsummary 1.5.1
torchvision 0.7.0+cu101
tqdm 4.66.1
typing_extensions 4.7.1
urllib3 2.0.7
wcwidth 0.2.9
Werkzeug 2.2.3
wheel 0.38.4
zipp 3.15.0

问题描述

想用GPU运行张量,一开始遇到输入张量与其他张量不在同一设备的错误,错误定位到F.linear(input, self.weight, self.bias)。

随后做了两处修改:

  • 自定义Linear层中,将self.weight = Parameter(torch.Tensor(out_features, in_features))改为self.weight = Parameter(torch.Tensor(out_features, in_features).to('cuda')),将self.bias = Parameter(torch.Tensor(out_features))改为self.bias = Parameter(torch.Tensor(out_features).to('cuda'));
  • SlotAttention模块中,将slots = torch.normal(mu, sigma)改为slots = torch.normal(mu, sigma).to('cuda')。

修改后出现新错误:RuntimeError: CUDA error: an illegal memory access was encountered。尝试将输入转为input = input.to('cuda'),错误依然存在。

修改后的代码

自定义Linear层

class Linear(Module):
    r"""Applies a linear transformation to the incoming data: :math:`y = xA^T + b`

    Args:
        in_features: size of each input sample
        out_features: size of each output sample
        bias: If set to ``False``, the layer will not learn an additive bias.
            Default: ``True``

    Shape:
        - Input: :math:`(N, *, H_{in})` where :math:`*` means any number of
          additional dimensions and :math:`H_{in} = \text{in\_features}`
        - Output: :math:`(N, *, H_{out})` where all but the last dimension
          are the same shape as the input and :math:`H_{out} = \text{out\_features}`.

    Attributes:
        weight: the learnable weights of the module of shape
            :math:`(\text{out\_features}, \text{in\_features})`. The values are
            initialized from :math:`\mathcal{U}(-\sqrt{k}, \sqrt{k})`, where
            :math:`k = \frac{1}{\text{in\_features}}`
        bias:   the learnable bias of the module of shape :math:`(\text{out\_features})`.
                If :attr:`bias` is ``True``, the values are initialized from
                :math:`\mathcal{U}(-\sqrt{k}, \sqrt{k})` where
                :math:`k = \frac{1}{\text{in\_features}}`

    Examples::

        >>> m = nn.Linear(20, 30)
        >>> input = torch.randn(128, 20)
        >>> output = m(input)
        >>> print(output.size())
        torch.Size([128, 30])
    """
    __constants__ = ['in_features', 'out_features']
    in_features: int
    out_features: int
    weight: Tensor

    def __init__(self, in_features: int, out_features: int, bias: bool = True) -> None:
        super(Linear, self).__init__()
        self.in_features = in_features
        self.out_features = out_features
        self.weight = Parameter(torch.Tensor(out_features, in_features).to('cuda'))
        if bias:
            self.bias = Parameter(torch.Tensor(out_features).to('cuda'))
        else:
            self.register_parameter('bias', None)
        self.reset_parameters()

    def reset_parameters(self) -> None:
        init.kaiming_uniform_(self.weight, a=math.sqrt(5))
        if self.bias is not None:
            fan_in, _ = init._calculate_fan_in_and_fan_out(self.weight)
            bound = 1 / math.sqrt(fan_in)
            init.uniform_(self.bias, -bound, bound)

    def forward(self, input: Tensor) -> Tensor:
        print("input.device", input.device)
        print("self.weight.device", self.weight.device)
        print("self.bias.device", self.bias.device)
        return F.linear(input, self.weight, self.bias)

    def extra_repr(self) -> str:
        return 'in_features={}, out_features={}, bias={}'.format(
            self.in_features, self.out_features, self.bias is not None
        )

SlotAttention模块

class SlotAttention(nn.Module):
    """
    Implementation from https://github.com/lucidrains/slot-attention by lucidrains.
    """
    def __init__(self, num_slots, dim, iters=3, eps=1e-8, hidden_dim=128):
        super().__init__()
        self.num_slots = num_slots
        self.iters = iters
        self.eps = eps
        self.scale = dim ** -0.5

        self.slots_mu = nn.Parameter(torch.randn(1, 1, dim))
        # self.slots_log_sigma = nn.Parameter(torch.randn(1, 1, dim))
        self.slots_log_sigma = nn.Parameter(torch.randn(1, 1, dim)).abs()

        self.project_q = nn.Linear(dim, dim)
        self.project_k = nn.Linear(dim, dim)
        self.project_v = nn.Linear(dim, dim)

        self.gru = nn.GRUCell(dim, dim)

        hidden_dim = max(dim, hidden_dim)

        self.mlp = nn.Sequential(
            nn.Linear(dim, hidden_dim),
            nn.ReLU(inplace=True),
            nn.Linear(hidden_dim, dim)
        )

        self.norm_inputs = nn.LayerNorm(dim, eps=1e-05)
        self.norm_slots = nn.LayerNorm(dim, eps=1e-05)
        self.norm_mlp = nn.LayerNorm(dim, eps=1e-05)

        # dummy initialisation
        self.attn = 0

    def forward(self, inputs, num_slots=None):
        b, n, d = inputs.shape
        n_s = num_slots if num_slots is not None else self.num_slots

        mu = self.slots_mu.expand(b, n_s, -1)
        sigma = self.slots_log_sigma.expand(b, n_s, -1)
        slots = torch.normal(mu, sigma).to('cuda')
        inputs = self.norm_inputs(inputs)
        k, v = self.project_k(inputs), self.project_v(inputs)

        for _ in range(self.iters):
            slots_prev = slots

            slots = self.norm_slots(slots)
            q = self.project_q(slots)

            dots = torch.einsum('bid,bjd->bij', q, k) * self.scale
            attn = dots.softmax(dim=1) + self.eps
            attn = attn / attn.sum(dim=-1, keepdim=True)

            updates = torch.einsum('bjd,bij->bid', v, attn)

            slots = self.gru(
                updates.reshape(-1, d),
                slots_prev.reshape(-1, d)
            )

            slots = slots.reshape(b, -1, d)
            slots = slots + self.mlp(self.norm_mlp(slots))

        self.attn = attn

        return slots

求助

有没有人遇到过类似问题?希望能得到帮助。


内容的提问来源于stack exchange,提问作者shuailei

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.07 06:05:57