You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

单隐藏层神经网络无法完成多类别预测的问题排查与算法建议

多类别分类模型问题求助

我正在搭建一个用于多类别分类任务的神经网络模型,因需要提取与自变量数量一致的权重用于其他任务,仅设置1层隐藏层。但模型表现不佳,混淆矩阵显示无法预测全部类别(仅能预测score为1和5的类别)。

数据情况

自变量为t1-t6,取值范围为[-∞, +∞];因变量为score,取值范围为[1-5],数据结构如下:

reviewIdt1t2t3t4t5t6score
01-3000001
0200380005
030900002
........................

模型代码

num_classes = 5
model_relu = Sequential()
model_relu.add(Dense(1, input_shape=(X.shape[1],), activation='relu')) # input shape is (features,). 1 hidden layer with 6 neurons
model_relu.add(Dense(num_classes, activation='softmax'))
model_relu.summary()
# compile the model
model_relu.compile(optimizer='rmsprop', 
              loss='categorical_crossentropy', # this is different instead of binary_crossentropy (for regular classification)
              metrics=['accuracy'])
# early stopping callback
# This callback will stop the training when there is no improvement in  
# the validation loss for 10 consecutive epochs.  
es = keras.callbacks.EarlyStopping(monitor='val_loss', 
                                   mode='min',
                                   patience=10, 
                                   restore_best_weights=True) # important - otherwise you just return the last weigths...

# now we just update our model fit call
history = model_relu.fit(X,
                    dummy_y,
                    callbacks=[es],
                    epochs=8000000, # you can set this to a big number!
                    batch_size=10000,
                    shuffle=True,
                    validation_split=0.2,
                    verbose=1)
preds = model_relu.predict(X) # see how the model did!
print(preds[0]) # i'm spreading that prediction across nodes and they sum to 1
print(np.sum(preds[0])) # sum it up! Should be 1
# confusion matrix
matrix = confusion_matrix(dummy_y.argmax(axis=1), preds.argmax(axis=1))
matrix

模型结果

混淆矩阵

array([[ 32141,      0,      0,      0, 114766],
       [  3500,      0,      0,      0,  19625],
       [  2541,      0,      0,      0,  27617],
       [  1657,      0,      0,      0,  42156],
       [ 10333,      0,      0,      0, 724463]])

模型评估结果

precision    recall  f1-score   support

           0       0.64      0.22      0.33    146907
           1       0.00      0.00      0.00     23125
           2       0.00      0.00      0.00     30158
           3       0.00      0.00      0.00     43813
           4       0.78      0.99      0.87    734796

    accuracy                           0.77    978799
   macro avg       0.28      0.24      0.24    978799
weighted avg       0.68      0.77      0.70    978799

请问我的代码是否存在错误?或是否有更合适的算法推荐?


问题分析与解决方案

代码中的核心错误

  1. 隐藏层神经元数量与注释不符
    代码注释标注“1 hidden layer with 6 neurons”,但实际用Dense(1)只设置了1个神经元。这直接导致模型特征提取能力严重不足,无法学习到中间类别(score=2、3、4)的区分模式,只能捕捉最极端的类别特征。

  2. 未做数据标准化
    自变量取值范围跨度极大(如t3=38、t1=-3),ReLU激活函数会直接屏蔽负输入的信息,大数值特征会主导模型学习,小数值特征的贡献被完全忽略,进一步恶化中间类别的识别效果。

  3. 未处理类别不平衡
    从评估结果的support列可见,score=5的样本量是其他类别的数倍甚至数十倍,模型会自然倾向于预测占比最高的类别,直接导致中间类别无法被识别。

针对性改进措施

修正模型结构

将隐藏层神经元数量改为6(匹配自变量数量,满足你提取对应权重的需求):

model_relu.add(Dense(6, input_shape=(X.shape[1],), activation='relu'))

数据预处理

对自变量做标准化,消除量纲影响:

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
# 后续使用X_scaled训练模型

处理类别不平衡

  • 类别权重调整:给少数类别更高的训练权重,平衡模型对不同类别的关注度:
    from sklearn.utils.class_weight import compute_class_weight
    class_weights = compute_class_weight('balanced', classes=np.unique(dummy_y.argmax(axis=1)), y=dummy_y.argmax(axis=1))
    class_weight_dict = dict(enumerate(class_weights))
    # 在fit中加入class_weight参数
    history = model_relu.fit(X_scaled, dummy_y, callbacks=[es], epochs=8000000, batch_size=10000, shuffle=True, validation_split=0.2, verbose=1, class_weight=class_weight_dict)
    
  • 采样调整:对少数类别做过采样,或对多数类别做欠采样,直接平衡样本分布。

训练细节优化

  • 降低batch_size:当前10000的batch_size过大,可能导致模型更新不充分,可尝试调整为512或1024。
  • 更换优化器:Adam优化器在多数分类任务中表现更稳定,可替换原有的rmsprop。

替代算法推荐

如果单隐藏层神经网络效果仍未达标,可尝试以下算法:

  • 带类别权重的逻辑回归:简单易解释,适合特征与类别线性相关的场景,同样可以提取特征权重。
  • 随机森林/梯度提升树:树模型对类别不平衡和非线性关系的处理能力更强,无需复杂预处理,还能输出特征重要性。
  • 支持向量机(SVM):在高维特征或样本量适中的场景下表现良好,可通过核函数捕捉非线性模式。

内容的提问来源于stack exchange,提问作者Dewani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 08:30:39