采用MealPy库中灰狼优化器(GWO)替代Adam优化器优化MLP时持续输出0.5准确率的问题排查求助
Let's break down why you're stuck at 0.5 accuracy and fix this step by step:
1. Critical Mistake in Binary Classification Prediction Handling
Your final layer uses sigmoid activation (correct for binary classification), but you're using np.argmax(yhat, axis=-1) to generate predictions—this is designed for multi-class tasks, not binary classification.
For a single-output sigmoid layer, each prediction is a scalar between 0 and 1. Running np.argmax on a scalar will always return 0, meaning all your predictions default to class 0. If your dataset is balanced (50% positive/negative samples), this leads directly to the 0.5 accuracy you're seeing.
Fix this in two key places:
- In the
fitness_functionmethod:yhat = self.model.predict(self.X_train) # Replace argmax with thresholding for binary classification yhat = (yhat >= 0.5).astype('int') # Convert probabilities to 0/1 labels acc = accuracy_score(self.y_train, yhat) - In the
predictionmethod (if you use it for test set evaluation later):yhat = self.model.predict(x_data) yhat = (yhat >= 0.5).astype('int') return yhat
2. Incorrect n_dims in Problem Definition
You hardcoded "n_dims": 35 in your problem configuration, but the actual number of parameters (dimensions) your optimizer needs to optimize is self.problem_size—the total count of weights and biases in your MLP. This means the optimizer is only updating 35 parameters instead of the full set, leaving most of your model's weights stuck at random initial values.
Fix:
Update your create_problem method to use the correct dimension count:
def create_problem(self): self.problem = { "fit_func": self.fitness_function, "lb": [-1, ] * self.problem_size, "ub": [1, ] * self.problem_size, "minmax": "max", "save_population": False, "n_dims": self.problem_size, # Match the actual number of model parameters }
3. Potential Gradient Vanishing with Stacked Sigmoid Layers
Stacking multiple sigmoid activation layers can cause gradient vanishing, making it extremely hard for any optimizer (evolutionary or gradient-based) to learn meaningful patterns. While this isn't the direct cause of your 0.5 accuracy, it will limit performance once you fix the above issues.
Suggestion:
Replace hidden layer activations with ReLU (or LeakyReLU) for better gradient flow:
model.add(Dense(self.n_hidden_nodes[0], input_dim=self.n_inputs, activation='relu')) model.add(Dense(self.n_hidden_nodes[1], activation='relu')) model.add(Dense(self.n_hidden_nodes[2], activation='relu'))
4. Verify Dataset Splitting and Preprocessing
Double-check that your dataset is properly split (no data leakage between train/test sets) and normalized. Neural networks, especially when paired with evolutionary optimizers, perform much better with scaled features (e.g., standardization to mean 0/std 1 or min-max scaling to [0,1]).
Quick sanity check:
Print the class distribution of self.y_train and self.y_test to confirm they aren't artificially skewed (though the first fix should resolve the 0.5 issue even with balanced data).
5. Adjust Optimizer Hyperparameters
Once the core fixes are applied, you may need to tweak MealPy's optimizer settings to get good performance:
- Increase
pop_size(try 50-100 instead of small values—evolutionary optimizers need enough diversity to explore the parameter space) - Extend the number of epochs (evolutionary methods often require more iterations than gradient-based optimizers like Adam)
- Narrow the search bounds (
lb/ub) if needed—starting with [-0.5, 0.5] instead of [-1,1] can help the optimizer converge faster.
After applying these changes, your model should start learning meaningful patterns and move beyond the 0.5 accuracy plateau. Test with GWO first, then revisit other optimizers to compare performance.
内容的提问来源于stack exchange,提问作者Warmn Faiq

