You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

M1 Mac使用MPS运行LSTM时loss.back()报错(CPU运行正常)

M1 MacBook Pro上MPS设备运行LSTM时loss.backward()报错问题

环境信息

  • Python版本:3.10.6
  • PyTorch版本:1.12.1(nightly构建)
  • 运行设备:M1 MacBook Pro

问题描述

LSTM网络代码在CPU上可正常运行,但切换到MPS设备时,执行loss.backward()会抛出如下错误:

RuntimeError: Expected a proper Tensor but got None (or an undefined Tensor in C++) for argument #0 'grad_y'

推测是梯度相关问题,但不清楚为何CPU上运行正常,也不知道该如何调试。

相关代码

模型定义

import torch
import torch.nn as nn

class ShallowRegressionLSTM(nn.Module):
    def __init__(self, num_sensors, hidden_units, num_layers=1, out_features=1):
        super(ShallowRegressionLSTM, self).__init__()
        self.num_sensors = num_sensors  # 特征数量
        self.hidden_units = hidden_units
        self.num_layers = num_layers
        self.out_features = out_features

        self.lstm = nn.LSTM(
            input_size=num_sensors,
            hidden_size=hidden_units,
            batch_first=True,
            num_layers=self.num_layers
        )

        self.linear = nn.Linear(in_features=self.hidden_units, out_features=self.out_features)

    def forward(self, x):
        batch_size = x.shape[0]
        h0 = torch.zeros(self.num_layers, batch_size, self.hidden_units, requires_grad=True, device=device)
        c0 = torch.zeros(self.num_layers, batch_size, self.hidden_units, requires_grad=True, device=device)
        
        _, (hn, _) = self.lstm(x, (h0, c0))  # 输出格式:output, (h_n, c_n)
        out = self.linear(hn[-1]).flatten()  # hn的第一维度是层数,取最后一层的隐藏状态
        return out

训练函数

def train_model(data_loader, model, loss_function, optimizer):
    num_batches = len(data_loader)
    total_loss = 0
    model.train()

    for X, y in data_loader:
        X = X.to(device)
        y = y.to(device)    
        output = model(X)
        
        loss = loss_function(output, y)
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

        total_loss += loss.item()

    avg_loss = total_loss / num_batches
    print(f"Train loss: {avg_loss}")
    train_loss.append(avg_loss)

模型初始化与训练循环

model = ShallowRegressionLSTM(num_sensors=len(features), hidden_units=num_hidden_units, num_layers=num_layers).to(device)
# loss_function = nn.MSELoss().to(device)
loss_function = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=learning_rate)
for epoch in range(epochs):
    train_model(train_loader, model, loss_function, optimizer=optimizer)

已尝试的解决方法

  • 将损失函数部署在GPU或CPU上
  • 使用retain_grad()方法
  • 在计算损失前将模型输出和标签移至CPU

内容的提问来源于stack exchange,提问作者Samed Ali

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 21:01:11