You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

输出层Softmax效果劣于ReLU?PyTorch二分类模型问题排查

问题描述

我有一个二类数据集,使用sklearn的SGDClassifier可得到测试集的完美混淆矩阵:

array([[1081, 0],
       [0, 982]])

我希望用PyTorch复现该结果,且不采用单输出的二分类方式,模型代码如下:

import torch.nn as nn

class TestClassificator(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.input_layer  = nn.Linear(12, 128)
        self.hidden_layer  = nn.Linear(128, 128)
        self.output_layer = nn.Linear(128, 2)
        self.hidden_activation = nn.ReLU()
        self.output_activation = nn.Softmax(dim=1)

    def forward(self, inputs):
        #
        x = self.input_layer(inputs)
        x = self.hidden_activation(x)
        #
        x = self.hidden_layer(x)
        x = self.output_activation(x)
        #
        x = self.output_layer(x)
        return x

我想使用Softmax使输出和为1,但此时混淆矩阵为:

from sklearn.metrics import confusion_matrix
from pandas import DataFrame

confusion_matrix(DataFrame(labels).values.argmax(axis=1), DataFrame(y_pred_list).values.argmax(axis=1))

array([[1072,    9],
       [0,  982]])

若输出层改用ReLU则能再次得到完美结果。请问我哪里出错了?为何Softmax的效果更差?


问题分析与解决

你的核心问题出在Softmax的位置错误,同时模型的特征传递逻辑也有疏漏:

  1. Softmax位置完全错误
    你将Softmax放在了隐藏层与输出层之间,这会严重破坏特征学习:

    • Softmax会把隐藏层的输出压缩到(0,1)区间且强制和为1,这会丢失隐藏层输出的特征幅度信息,后续的线性输出层无法基于原始的特征差异进行有效拟合,模型的表达能力被大幅限制。
    • 正确的做法是将Softmax放在输出层之后:先让输出层输出未归一化的logits,再用Softmax将其转换为概率分布。
  2. 隐藏层缺少激活函数
    你的forward函数中,隐藏层计算后直接接了Softmax,漏掉了定义好的hidden_activation(ReLU)。这会让隐藏层退化为线性变换,无法学习复杂的非线性特征,进一步降低模型性能。

  3. 二分类场景下的Softmax冗余性
    对于二类任务,用Softmax(dim=1)虽然能得到和为1的输出,但其实用Sigmoid更高效——只需预测一类的概率,另一类概率可通过1减去该值得到,还能避免Softmax带来的潜在数值稳定性问题。不过如果坚持使用Softmax,只要位置正确即可。

  4. ReLU能完美拟合的原因
    当输出层改用ReLU时,模型输出的是未归一化的logits,模型可以自由学习到足够大的特征幅度差异,让argmax能准确区分两类样本,因此能复现完美混淆矩阵。

修正后的模型代码

import torch.nn as nn

class TestClassificator(nn.Module):
    def __init__(self) -> None:
        super().__init__()
        self.input_layer  = nn.Linear(12, 128)
        self.hidden_layer  = nn.Linear(128, 128)
        self.output_layer = nn.Linear(128, 2)
        self.hidden_activation = nn.ReLU()
        self.output_activation = nn.Softmax(dim=1)

    def forward(self, inputs):
        x = self.input_layer(inputs)
        x = self.hidden_activation(x)
        
        x = self.hidden_layer(x)
        x = self.hidden_activation(x)  # 补上隐藏层的ReLU激活
        
        x = self.output_layer(x)
        x = self.output_activation(x)  # 将Softmax移到输出层之后
        return x

内容的提问来源于stack exchange,提问作者R. Key

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 21:05:19