基于MLP的二分类欠拟合问题求助——CTU-13流量分类
Hey, let's break down your problem step by step. You're working on a supervised anomaly detection MLP for benign/malicious traffic classification using the CTU-13 dataset, with current metrics: Accuracy 0.923, Precision 0.999, Recall 0.833, F1 0.909. You suspect underfitting after trying to add neurons and layers without improvement. Here are actionable suggestions to fix this:
First, Validate If It's Actually Underfitting
Underfitting typically shows up as high loss on both training and validation sets, with minimal gap between the two. If your training accuracy is way higher than validation accuracy, you might be dealing with overfitting instead. Double-check your training curves (loss/accuracy over epochs) to confirm—sharing those would help us be more precise.
1. Fix Data Preprocessing (Often Overlooked!)
MLPs are extremely sensitive to feature scaling and quality:
- Standardize/Normalize Features: Your 15-dimensional input features likely have varying scales. Use
StandardScalerorMinMaxScalerto bring all features into the same range. Without this, ReLU neurons can easily saturate, slowing down learning and mimicking underfitting. - Boost Feature Engineering: CTU-13's raw features might be redundant or lack meaningful patterns. Try extracting statistical features (like packet size mean/variance, connection duration) or using PCA to reduce noise while preserving key information. Better features = easier model learning.
- Address Class Imbalance (Even If Subtle): While your dataset is roughly balanced (169k benign vs 143k malicious), malicious traffic might have more scattered features. Try:
- Over-sampling with SMOTE for the minority class
- Adding
class_weightinmodel.compileto assign higher weight to malicious samples (e.g.,class_weight={0:1, 1:1.2}if 1 represents malicious traffic)
2. Tune Model Architecture & Training Parameters
You tried adding layers/neurons, but let's adjust the approach:
- Replace ReLU with LeakyReLU/ELU: Plain ReLU can create "dead neurons" that stop learning. Swap in LeakyReLU to preserve gradient flow:
from tensorflow.keras.layers import LeakyReLU model.add(Dense(1024, input_dim=15)) model.add(LeakyReLU(alpha=0.1)) - Adjust Dropout & Add Batch Normalization: Your 0.5 dropout rate might be too aggressive, especially in shallow layers. Drop it to 0.2-0.3, and pair layers with BatchNormalization to stabilize training:
from tensorflow.keras.layers import BatchNormalization model.add(Dense(1024, input_dim=15)) model.add(BatchNormalization()) model.add(LeakyReLU(alpha=0.1)) model.add(Dropout(0.3)) - Optimize Learning Rate: Your Adam lr=0.0001 is probably too small, leading to slow convergence. Bump it to 0.001, and add a learning rate scheduler to fine-tune later:
from tensorflow.keras.callbacks import ReduceLROnPlateau callbacks = [ EarlyStopping(monitor='val_loss', patience=8, restore_best_weights=True), ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3, min_lr=1e-6) ] - Adjust Batch Size & Epochs: A batch size of 50 might introduce too much noise. Try 128 or 256, and extend
patiencein EarlyStopping to 8-10 to give the model more time to converge.
3. Fix the Recall Gap (High Precision, Low Recall)
Your metrics show the model is great at avoiding false positives but misses many malicious samples. This isn't just underfitting—try:
- Analyzing misclassified malicious samples to find feature patterns the model is missing
- Using a confusion matrix to pinpoint exactly which samples are being mislabeled
- Trying a focal loss function instead of binary crossentropy, which prioritizes hard-to-classify samples
Example Adjusted Model
Here's a revised version of your code incorporating these changes:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Dropout, BatchNormalization, LeakyReLU from tensorflow.keras import optimizers from tensorflow.keras.callbacks import EarlyStopping, ReduceLROnPlateau def MLP_model(): model = Sequential() # Input layer with BatchNorm and LeakyReLU model.add(Dense(1024, input_dim=15)) model.add(BatchNormalization()) model.add(LeakyReLU(alpha=0.1)) model.add(Dropout(0.3)) # Hidden layers model.add(Dense(512)) model.add(BatchNormalization()) model.add(LeakyReLU(alpha=0.1)) model.add(Dense(256)) model.add(BatchNormalization()) model.add(LeakyReLU(alpha=0.1)) model.add(Dropout(0.3)) model.add(Dense(128)) model.add(BatchNormalization()) model.add(LeakyReLU(alpha=0.1)) # Output layer model.add(Dense(1, activation='sigmoid')) # Optimizer with adjusted learning rate adam = optimizers.Adam(lr=0.001, beta_1=0.9, beta_2=0.999, epsilon=1e-07, decay=0.0, amsgrad=False) model.compile(optimizer=adam, loss='binary_crossentropy', metrics=['accuracy']) return model model = MLP_model() # Enhanced callbacks callbacks = [ EarlyStopping(monitor='val_loss', patience=8, restore_best_weights=True), ReduceLROnPlateau(monitor='val_loss', factor=0.5, patience=3, min_lr=1e-6) ] # Add class weight to prioritize malicious samples class_weight = {0: 1.0, 1: 1.2} hist = model.fit(Xtrain, Ytrain, epochs=100, batch_size=128, validation_split=0.20, callbacks=callbacks, verbose=1, class_weight=class_weight)
Give these changes a shot, and if you can share your training curves or more details about your feature set, we can refine this further.
内容的提问来源于stack exchange,提问作者benynugraha

