PyTorch训练MLP遇mat1与mat2 dtype不一致错误求助
问题解决方法
错误根源
mat1 and mat2 must have the same dtype 错误的核心是输入张量与模型参数的数据类型不匹配,同时代码还存在拼写错误和梯度清零逻辑错误。
修正步骤及完整代码
1. 统一数据类型
确保输入张量的 dtype 与模型参数完全一致,模型默认使用torch.float32,转换numpy数组时需显式指定:
# 转换numpy数组为torch张量,指定dtype和设备 x = torch.from_numpy(x).to(dtype=dtype, device=device) y = torch.from_numpy(y).to(dtype=dtype, device=device) # 将模型同步到对应dtype和设备 model = model.to(dtype=dtype, device=device)
2. 修正拼写与梯度清零逻辑
- 修正
learning_rat的拼写错误为learning_rate - 替换错误的
param = None,改用param.grad.zero_()清零梯度,避免梯度累积
修正后的完整代码
def mlp_gradient_descent(x, y, model, eta=1e-6, nb_iter=30000): loss_descent = [] dtype = torch.float32 device = torch.device("cpu") # 转换输入并统一数据类型与设备 x = torch.from_numpy(x).to(dtype=dtype, device=device) y = torch.from_numpy(y).to(dtype=dtype, device=device) model = model.to(dtype=dtype, device=device) params = model.parameters() learning_rate = eta for t in range(nb_iter): y_pred = model(x) loss = (y_pred - y).pow(2).sum() if t % 100 == 99: print(t, loss.item()) loss_descent.append([t, loss.item()]) loss.backward() with torch.no_grad(): # 手动更新参数 for param in params: param -= learning_rate * param.grad # 清零梯度,避免累积 param.grad.zero_() return loss_descent
额外说明
- 若你的numpy数组是
float64类型,可将dtype改为torch.float64,同时同步模型的dtype即可 - 所有张量与模型必须处于同一设备(CPU/GPU),否则也会引发类型不兼容错误
内容的提问来源于stack exchange,提问作者Bendaoud Simou
相关产品推荐
相关产品推荐

