You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在PyTorch中复现TensorFlow Sequential多分类模型遇维度错误求帮助

问题描述

我有一个用于多分类任务的TensorFlow浅层次Sequential模型:

model = tf.keras.models.Sequential([  
    tf.keras.layers.Dense(64, input_shape = (384,), activation = "relu"),
    tf.keras.layers.Dense(36, activation="softmax", use_bias = False)
])    
        
model.summary()  

# 输出:
Model: "sequential_5"
_________________________________________________________________
 Layer (type)                Output Shape              Param #   
=================================================================
 dense_13 (Dense)            (None, 64)                24640     
                                                                 
 dense_14 (Dense)            (None, 36)                2304 

我希望在PyTorch中复现该模型,并在forward函数中使用softmax。我的基础模型可以运行但准确率很低,代码如下:

# base model - it works but gives v low accuracy. need to add more layers.

class LR(torch.nn.Module):

    def __init__(self, n_features):
        super(LR, self).__init__()
        self.lr = torch.nn.Linear(n_features, 36)
        nn.ReLU()

        
    def forward(self, x):
        out = torch.softmax(self.lr(x),dim = 1,dtype=None)
        return out

# parameters
n_features = 384
n_classes = 36
optim = torch.optim.SGD(model.parameters(), lr=0.1)
criterion = torch.nn.CrossEntropyLoss()

accuracy_fn = Accuracy(task="multiclass",  num_classes=36)

我尝试修改模型类后出现了错误,修改后的代码及错误信息如下:

class LR(torch.nn.Module):

    def __init__(self, n_features):
        super(LR, self).__init__()
        self.lr = torch.nn.Linear(n_features, 64)        
        self.lr = nn.ReLU()
        self.lr = torch.nn.Linear(64, 36)
        
    def forward(self, x):
        out = torch.softmax(self.lr(x),dim = 1,dtype=None)
        return out

epochs = 15

def train(model, optim, criterion, x, y, epochs=epochs):
    for e in range(1, epochs + 1):
        optim.zero_grad()
        out = model(x)
        loss = criterion(out, y)
        loss.backward()
        optim.step()
        print(f"Loss at epoch {e}: {loss.data}")

    return model

model = LR(n_features)

model = train(model, optim, criterion, X_train, y_train)

错误信息:

RuntimeError                              Traceback (most recent call last)
Cell In[175], line 18
     12  #       acc = accuracy_fn(out, y)  
     13  #       acc.backward()
     14  #       optim.step()
     15  #       print(f"Acc at epoch {e}: {acc.data}")
     16     return model
---> 18 model = train(model, optim, criterion, X_train, y_train)

Cell In[175], line 6, in train(model, optim, criterion, x, y, epochs)
      4 for e in range(1, epochs + 1):
      5     optim.zero_grad()
----> 6     out = model(x)
      7     loss = criterion(out, y)
      8     loss.backward()

File /usr/local/lib/python3.10/site-packages/torch/nn/modules/module.py:1501, in Module._call_impl(self, *args, **kwargs)
   1496 # If we don't have any hooks, we want to skip the rest of the logic in
   1497 # this function, and just call forward.
   1498 if not (self._backward_hooks or self._backward_pre_hooks or self._forward_hooks or self._forward_pre_hooks
   1499         or _global_backward_pre_hooks or _global_backward_hooks
   1500         or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501     return forward_call(*args, **kwargs)
   1502 # Do not call functions when jit is used
   1503 full_backward_hooks, non_full_backward_hooks = [], []

Cell In[172], line 10, in LR.forward(self, x)
      9 def forward(self, x):
---> 10     out = torch.softmax(self.lr(x),dim = 1,dtype=None)
     11     return out

File /usr/local/lib/python3.10/site-packages/torch/nn/modules/module.py:1501, in Module._call_impl(self, *args, **kwargs)
   1496 # If we don't have any hooks, we want to skip the rest of the logic in
   1497 # this function, and just call forward.
   1498 if not (self._backward_hooks or self._backward_pre_hooks or self._forward_hooks or self._forward_pre_hooks
   1499         or _global_backward_pre_hooks or _global_backward_hooks
   1500         or _global_forward_hooks or _global_forward_pre_hooks):
-> 1501     return forward_call(*args, **kwargs)
   1502 # Do not call functions when jit is used
   1503 full_backward_hooks, non_full_backward_hooks = [], []

File /usr/local/lib/python3.10/site-packages/torch/nn/modules/linear.py:114, in Linear.forward(self, input)
    113 def forward(self, input: Tensor) -> Tensor:
--> 114     return F.linear(input, self.weight, self.bias)

RuntimeError: mat1 and mat2 shapes cannot be multiplied (155648x384 and 64x36)

训练数据形状:

X_train.shape  # torch.Size([155648, 384]) 
y_train.shape  # torch.Size([155648])
解决方案

错误原因分析

  1. 模型层覆盖:修改后的LR类中,self.lr被重复赋值,最终仅保留最后一层Linear(64,36),导致输入维度(155648,384)与该层权重维度(64,36)不匹配,引发形状错误。
  2. 激活层无效:原代码中nn.ReLU()仅创建实例但未绑定到模型属性,也未在forward中调用,完全不起作用。
  3. 损失函数误用:PyTorch的CrossEntropyLoss内部已包含softmax计算,手动在forward中添加softmax会导致损失计算异常,影响训练效果。
  4. 初始化顺序错误:原代码先定义优化器,但此时模型未实例化,会引发参数获取失败的问题。

修正后的代码

1. 正确的模型类实现

import torch
import torch.nn as nn
from torchmetrics import Accuracy

class LR(nn.Module):
    def __init__(self, n_features):
        super(LR, self).__init__()
        # 对应TensorFlow第一层Dense(64, relu)
        self.dense1 = nn.Linear(n_features, 64)
        self.relu = nn.ReLU()
        # 对应TensorFlow第二层Dense(36, softmax, use_bias=False)
        self.dense2 = nn.Linear(64, 36, bias=False)
        
    def forward(self, x):
        x = self.dense1(x)
        x = self.relu(x)
        x = self.dense2(x)
        # CrossEntropyLoss自带softmax,此处无需手动添加
        # 推理阶段如需输出概率,可调用torch.softmax(x, dim=1)
        return x

2. 修正后的训练流程与参数配置

# 参数配置
n_features = 384
n_classes = 36
epochs = 15

# 实例化模型、优化器、损失函数、准确率指标
model = LR(n_features)
optim = torch.optim.SGD(model.parameters(), lr=0.1)
criterion = nn.CrossEntropyLoss()
accuracy_fn = Accuracy(task="multiclass", num_classes=36)

def train(model, optim, criterion, accuracy_fn, X_train, y_train, epochs=epochs):
    model.train()  # 开启训练模式
    for e in range(1, epochs + 1):
        optim.zero_grad()
        # 前向传播
        outputs = model(X_train)
        # 计算损失
        loss = criterion(outputs, y_train)
        # 反向传播+优化
        loss.backward()
        optim.step()
        # 计算准确率
        acc = accuracy_fn(outputs.argmax(dim=1), y_train)
        # 打印训练日志
        print(f"Epoch {e}/{epochs} | Loss: {loss.item():.4f} | Accuracy: {acc.item():.4f}")
    return model

# 启动训练
model = train(model, optim, criterion, accuracy_fn, X_train, y_train)

关键说明

  • 层命名规范:每个层使用独立属性名(如dense1、relu、dense2),避免覆盖。
  • 损失函数适配:CrossEntropyLoss要求输入为未经过softmax的logits,因此forward中无需手动添加softmax;若需输出类别概率,可在推理阶段对模型结果调用torch.softmax(outputs, dim=1)。
  • 训练模式:调用model.train()确保模型处于训练状态(对Dropout、BatchNorm等层生效,养成规范习惯)。
  • 准确率计算:通过outputs.argmax(dim=1)获取预测类别,再与真实标签计算准确率。

内容的提问来源于stack exchange,提问作者Bluetail

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 10:34:59