Keras中Autoencoder超参数调优求助:训练loss恒定无变化
Hey there! Let's work through your problem step by step—since you're new to Python and Keras, I'll keep things clear and actionable. First, let's tackle implementing grid search for your autoencoder, then dive into other tuning strategies to fix that stuck loss value.
Keras models need a wrapper to work with scikit-learn's GridSearchCV. Since you're building an autoencoder (a regression task where we reconstruct input data), we'll use KerasRegressor. Here's a complete, adaptable code example:
import numpy as np from keras.models import Model from keras.layers import Input, Dense from keras.wrappers.scikit_learn import KerasRegressor from sklearn.model_selection import GridSearchCV # Fixed parameters you mentioned INPUT_SHAPE = X_train.shape[1] # Match your input layer size ENCODING_DIM = 32 # Keep this fixed as required # Define a model-building function (grid search needs this to test parameters) def build_autoencoder(optimizer="adam", learning_rate=0.001, hidden_units=64): # Input layer input_layer = Input(shape=(INPUT_SHAPE,)) # Encoding layer (fixed dimension) encoder = Dense(ENCODING_DIM, activation="relu")(input_layer) # Optional hidden layer (tunable parameter) hidden = Dense(hidden_units, activation="relu")(encoder) # Decoding layer (matches input shape) decoder = Dense(INPUT_SHAPE, activation="linear")(hidden) # Linear for continuous input (-6 to 6) model = Model(inputs=input_layer, outputs=decoder) # Handle optimizer with custom learning rate if optimizer == "adam": from keras.optimizers import Adam opt = Adam(learning_rate=learning_rate) elif optimizer == "sgd": from keras.optimizers import SGD opt = SGD(learning_rate=learning_rate, momentum=0.9) model.compile(optimizer=opt, loss="mse") # MSE is ideal for reconstruction tasks return model # Wrap the Keras model for scikit-learn autoencoder = KerasRegressor(build_fn=build_autoencoder, verbose=0) # Define your parameter grid to search param_grid = { "optimizer": ["adam", "sgd", "rmsprop"], "learning_rate": [0.001, 0.0001, 0.01], "hidden_units": [32, 64, 128], # Remove this if you don't have a hidden layer "epochs": [50, 100], "batch_size": [16, 32, 64] } # Run grid search with 3-fold cross-validation grid = GridSearchCV(estimator=autoencoder, param_grid=param_grid, cv=3, scoring="neg_mean_squared_error") grid_result = grid.fit(X_train, X_train) # Autoencoders use X_train as both input and target # Print results print(f"Best validation loss: {grid_result.best_score_:.4f} using parameters: {grid_result.best_params_}") for mean, stdev, param in zip(grid_result.cv_results_["mean_test_score"], grid_result.cv_results_["std_test_score"], grid_result.cv_results_["params"]): print(f"Mean score: {mean:.4f} (std: {stdev:.4f}) | Params: {param}")
Key Notes for This Code:
- Since it's an autoencoder, we pass
X_trainas both the input and target tofit(). - If your model doesn't have a hidden layer, remove the
hidden_unitsentry fromparam_gridand the corresponding layer inbuild_autoencoder(). - Use
verbose=1inKerasRegressorif you want to see training progress during grid search.
A loss that doesn't move at all means your model isn't learning—let's troubleshoot and tune:
Verify Data Preprocessing:
Double-check your scaled data: Are there any NaNs or outliers? Runnp.isnan(X_train).any()to confirm. Also, ensure you're using a loss function that fits your data: since your input is continuous (-6 to 6), MSE or MAE is correct—avoid classification losses likebinary_crossentropy.Adjust Model Architecture:
Even with fixed input/encoding dimensions, you can tweak:- Activation functions: Try
tanhinstead ofreluin the encoding layer (matches your input's negative range). - Batch normalization: Add
BatchNormalization()after the encoding/hidden layers to stabilize training:from keras.layers import BatchNormalization encoder = Dense(ENCODING_DIM, activation="relu")(input_layer) encoder = BatchNormalization()(encoder) - Weight initialization: Use
kernel_initializer="he_normal"in dense layers to give the model a better starting point.
- Activation functions: Try
Optimizer & Training Tweaks:
- Ensure
shuffle=Trueinmodel.fit()to prevent the model from memorizing data order. - Try adding weight decay to the optimizer (e.g.,
Adam(learning_rate=0.001, weight_decay=1e-5)) to prevent overfitting (though your issue is underfitting right now). - Test smaller batch sizes—sometimes larger batches can stall learning if the learning rate isn't tuned.
- Ensure
Simplify First:
Strip the model down to the basics (only input → encoding → decoding layers) and see if the loss moves. If it still doesn't, the problem is likely with your data or optimizer setup, not model complexity.
- Use
model.summary()to double-check layer dimensions match your expectations. - Start with small training runs (e.g., 10 epochs) to test changes faster.
- Print a sample of your input data to confirm scaling was applied correctly.
内容的提问来源于stack exchange,提问作者pynewb

