You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras多分类MLP分类器准确率低问题求助

Troubleshooting Your MLP Classifier Stuck Predicting Only the First Class

Hey there! Let's break down why your model is only predicting the first class and hovering around 33% accuracy (which is just random chance for 3 classes). I'll walk through the key issues and fixes step by step:

1. Critical: Mismatched Output Layer Size

The biggest red flag here is your output layer:

model.add(Dense(4, activation='softmax'))

You have 3 classes, but you're using 4 neurons in the softmax layer. When you run to_categorical(y_train) on 3 classes, it creates 3-dimensional one-hot vectors—but your model is outputting 4 values. This misalignment confuses the model's loss calculation, making it impossible to learn proper class distinctions.

Fix: Change the output layer to match your number of classes:

model.add(Dense(3, activation='softmax'))

2. Missing Feature Scaling

MLPs (especially when using SGD) are extremely sensitive to feature scaling. Your single feature likely has a wide range of values, which makes it hard for the optimizer to converge. Without scaling, the model can't learn meaningful patterns from the feature.

Fix: Add standardization or normalization to your features:

from sklearn.preprocessing import StandardScaler

# Fit scaler on training data only, then apply to both train and test
scaler = StandardScaler()
x_train = scaler.fit_transform(x_train)
x_test = scaler.transform(x_test)

3. Potential Class Imbalance

From your y_train preview, it looks like the first class might dominate the dataset. If one class has way more samples than the others, the model will naturally bias toward predicting that class to minimize loss.

Check your class distribution:

print("Training class counts:", np.sum(y_train, axis=0))
print("Test class counts:", np.sum(y_test, axis=0))

Fixes if imbalance exists:

  • Use class_weight='balanced' in your fit() call to let Keras weight minority classes more heavily:
    history = model.fit(x_train, y_train, epochs=100, batch_size=128, class_weight='balanced')
    
  • Or use resampling techniques (oversample minority classes, undersample majority class) with libraries like imblearn.

4. Prediction Post-Processing Issue

Your current code uses np.around(prediction) before taking argmax, which is incorrect. Softmax outputs probabilities (values between 0 and 1), and rounding them can create invalid one-hot vectors (e.g., multiple 1s or all 0s), leading to wrong class assignments.

Fix: Directly take the index of the highest probability:

prediction = model.predict(x_test)
y_predict_non_category = np.argmax(prediction, axis=1)  # No need for np.around()
y_test_non_category = np.argmax(y_test, axis=1)

conf_mat = confusion_matrix(y_test_non_category, y_predict_non_category)
print(conf_mat)

5. Minor Tuning Suggestions

  • Batch Size: If your dataset is small, batch_size=128 might be too large—try reducing it to 32 or 16 to give the optimizer more frequent updates.
  • Epochs: 100 epochs might not be enough for convergence. Try increasing to 200, and add validation_split=0.1 to monitor overfitting:
    history = model.fit(x_train, y_train, epochs=200, batch_size=32, validation_split=0.1, class_weight='balanced')
    
  • Activation Functions: While tanh works, ReLU is often more stable for MLPs. Try swapping hidden layer activations to relu and see if performance improves:
    model.add(Dense(64, activation='relu', input_dim=1))
    

Full Corrected Code Snippet

Here's how your code should look after applying the key fixes:

import keras
from keras.models import Sequential
from keras.layers import Dense, Dropout, Activation
from keras.optimizers import SGD
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.metrics import confusion_matrix

# Importing the dataset
dataset = pd.read_csv('StatusAll.csv')
X = dataset.iloc[:, 1:].values
y = dataset.iloc[:, 0:1].values

# Splitting the dataset
x_train, x_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)

# Feature scaling
scaler = StandardScaler()
x_train = scaler.fit_transform(x_train)
x_test = scaler.transform(x_test)

# Convert labels to categorical
from keras.utils import to_categorical
y_train = to_categorical(y_train)
y_test = to_categorical(y_test)

# Check class distribution
print("Training class counts:", np.sum(y_train, axis=0))
print("Test class counts:", np.sum(y_test, axis=0))

# Build model
model = Sequential()
model.add(Dense(64, activation='relu', input_dim=1))
model.add(Dropout(0.5))
model.add(Dense(64, activation='relu'))
model.add(Dropout(0.5))
model.add(Dense(3, activation='softmax'))  # Corrected to 3 classes

# Compile model
sgd = SGD(lr=0.01, decay=1e-6, momentum=0.9, nesterov=True)
model.compile(loss='categorical_crossentropy', optimizer=sgd, metrics=['accuracy'])

# Train model
history = model.fit(x_train, y_train, epochs=200, batch_size=32, validation_split=0.1, class_weight='balanced')

# Evaluate
score = model.evaluate(x_test, y_test, batch_size=32)
print('Test score:', score[0])
print('Test accuracy:', score[1])

# Predict and confusion matrix
prediction = model.predict(x_test)
y_test_non_category = np.argmax(y_test, axis=1)
y_predict_non_category = np.argmax(prediction, axis=1)
conf_mat = confusion_matrix(y_test_non_category, y_predict_non_category)
print(conf_mat)

Start by fixing the output layer size and adding feature scaling—those two changes alone should make a huge difference. Then check for class imbalance and adjust training parameters as needed.

内容的提问来源于stack exchange,提问作者To win

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:57:29