Keras多分类MLP分类器准确率低问题求助
Hey there! Let's break down why your model is only predicting the first class and hovering around 33% accuracy (which is just random chance for 3 classes). I'll walk through the key issues and fixes step by step:
1. Critical: Mismatched Output Layer Size
The biggest red flag here is your output layer:
model.add(Dense(4, activation='softmax'))
You have 3 classes, but you're using 4 neurons in the softmax layer. When you run to_categorical(y_train) on 3 classes, it creates 3-dimensional one-hot vectors—but your model is outputting 4 values. This misalignment confuses the model's loss calculation, making it impossible to learn proper class distinctions.
Fix: Change the output layer to match your number of classes:
model.add(Dense(3, activation='softmax'))
2. Missing Feature Scaling
MLPs (especially when using SGD) are extremely sensitive to feature scaling. Your single feature likely has a wide range of values, which makes it hard for the optimizer to converge. Without scaling, the model can't learn meaningful patterns from the feature.
Fix: Add standardization or normalization to your features:
from sklearn.preprocessing import StandardScaler # Fit scaler on training data only, then apply to both train and test scaler = StandardScaler() x_train = scaler.fit_transform(x_train) x_test = scaler.transform(x_test)
3. Potential Class Imbalance
From your y_train preview, it looks like the first class might dominate the dataset. If one class has way more samples than the others, the model will naturally bias toward predicting that class to minimize loss.
Check your class distribution:
print("Training class counts:", np.sum(y_train, axis=0)) print("Test class counts:", np.sum(y_test, axis=0))
Fixes if imbalance exists:
- Use
class_weight='balanced'in yourfit()call to let Keras weight minority classes more heavily:history = model.fit(x_train, y_train, epochs=100, batch_size=128, class_weight='balanced') - Or use resampling techniques (oversample minority classes, undersample majority class) with libraries like
imblearn.
4. Prediction Post-Processing Issue
Your current code uses np.around(prediction) before taking argmax, which is incorrect. Softmax outputs probabilities (values between 0 and 1), and rounding them can create invalid one-hot vectors (e.g., multiple 1s or all 0s), leading to wrong class assignments.
Fix: Directly take the index of the highest probability:
prediction = model.predict(x_test) y_predict_non_category = np.argmax(prediction, axis=1) # No need for np.around() y_test_non_category = np.argmax(y_test, axis=1) conf_mat = confusion_matrix(y_test_non_category, y_predict_non_category) print(conf_mat)
5. Minor Tuning Suggestions
- Batch Size: If your dataset is small,
batch_size=128might be too large—try reducing it to 32 or 16 to give the optimizer more frequent updates. - Epochs: 100 epochs might not be enough for convergence. Try increasing to 200, and add
validation_split=0.1to monitor overfitting:history = model.fit(x_train, y_train, epochs=200, batch_size=32, validation_split=0.1, class_weight='balanced') - Activation Functions: While
tanhworks, ReLU is often more stable for MLPs. Try swapping hidden layer activations toreluand see if performance improves:model.add(Dense(64, activation='relu', input_dim=1))
Full Corrected Code Snippet
Here's how your code should look after applying the key fixes:
import keras from keras.models import Sequential from keras.layers import Dense, Dropout, Activation from keras.optimizers import SGD import pandas as pd import numpy as np from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.metrics import confusion_matrix # Importing the dataset dataset = pd.read_csv('StatusAll.csv') X = dataset.iloc[:, 1:].values y = dataset.iloc[:, 0:1].values # Splitting the dataset x_train, x_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) # Feature scaling scaler = StandardScaler() x_train = scaler.fit_transform(x_train) x_test = scaler.transform(x_test) # Convert labels to categorical from keras.utils import to_categorical y_train = to_categorical(y_train) y_test = to_categorical(y_test) # Check class distribution print("Training class counts:", np.sum(y_train, axis=0)) print("Test class counts:", np.sum(y_test, axis=0)) # Build model model = Sequential() model.add(Dense(64, activation='relu', input_dim=1)) model.add(Dropout(0.5)) model.add(Dense(64, activation='relu')) model.add(Dropout(0.5)) model.add(Dense(3, activation='softmax')) # Corrected to 3 classes # Compile model sgd = SGD(lr=0.01, decay=1e-6, momentum=0.9, nesterov=True) model.compile(loss='categorical_crossentropy', optimizer=sgd, metrics=['accuracy']) # Train model history = model.fit(x_train, y_train, epochs=200, batch_size=32, validation_split=0.1, class_weight='balanced') # Evaluate score = model.evaluate(x_test, y_test, batch_size=32) print('Test score:', score[0]) print('Test accuracy:', score[1]) # Predict and confusion matrix prediction = model.predict(x_test) y_test_non_category = np.argmax(y_test, axis=1) y_predict_non_category = np.argmax(prediction, axis=1) conf_mat = confusion_matrix(y_test_non_category, y_predict_non_category) print(conf_mat)
Start by fixing the output layer size and adding feature scaling—those two changes alone should make a huge difference. Then check for class imbalance and adjust training parameters as needed.
内容的提问来源于stack exchange,提问作者To win

