You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PyTorch训练MLP遇mat1与mat2 dtype不一致错误求助

问题解决方法

错误根源

mat1 and mat2 must have the same dtype 错误的核心是输入张量与模型参数的数据类型不匹配,同时代码还存在拼写错误和梯度清零逻辑错误。

修正步骤及完整代码

1. 统一数据类型

确保输入张量的 dtype 与模型参数完全一致,模型默认使用torch.float32,转换numpy数组时需显式指定:

# 转换numpy数组为torch张量,指定dtype和设备
x = torch.from_numpy(x).to(dtype=dtype, device=device)
y = torch.from_numpy(y).to(dtype=dtype, device=device)
# 将模型同步到对应dtype和设备
model = model.to(dtype=dtype, device=device)

2. 修正拼写与梯度清零逻辑

  • 修正learning_rat的拼写错误为learning_rate
  • 替换错误的param = None,改用param.grad.zero_()清零梯度,避免梯度累积

修正后的完整代码

def mlp_gradient_descent(x, y, model, eta=1e-6, nb_iter=30000): 
    loss_descent = []
    dtype = torch.float32
    device = torch.device("cpu")
    
    # 转换输入并统一数据类型与设备
    x = torch.from_numpy(x).to(dtype=dtype, device=device)
    y = torch.from_numpy(y).to(dtype=dtype, device=device)
    model = model.to(dtype=dtype, device=device)
    
    params = model.parameters()
    learning_rate = eta
    
    for t in range(nb_iter):
        y_pred = model(x)
        loss = (y_pred - y).pow(2).sum()
        
        if t % 100 == 99:
            print(t, loss.item())
            loss_descent.append([t, loss.item()])
        
        loss.backward()
        
        with torch.no_grad():
            # 手动更新参数
            for param in params:
                param -= learning_rate * param.grad
                # 清零梯度,避免累积
                param.grad.zero_()
    
    return loss_descent

额外说明

  • 若你的numpy数组是float64类型,可将dtype改为torch.float64,同时同步模型的dtype即可
  • 所有张量与模型必须处于同一设备(CPU/GPU),否则也会引发类型不兼容错误

内容的提问来源于stack exchange,提问作者Bendaoud Simou

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 18:50:40