如何在Keras的TensorBoard中按模型分组直方图?
Hey there, let's tackle this histogram grouping issue you're hitting with TensorBoard during your MNIST classification HPARAM tuning. The core problem here is that you're writing all your model logs into the same root directory—TensorBoard merges all data under the same folder, which is why your weight/bias histograms are showing up as a single continuous sequence (W1-W6, W7-W12, etc.) instead of being grouped per model.
The Right Approach: Independent Log Directories per Run
TensorBoard relies on separate subdirectories to distinguish different training runs. This is also the standard way to work with the HPARAM plugin, since it ties each set of hyperparameters to its own log folder. Adding tf.name_scope won't fix this because it only adds a prefix within the same log context—you need full directory separation.
Step-by-Step Code Modifications
Here's how to adjust your code to get proper grouping:
Generate a unique log subdirectory for each hyperparameter trial
We'll create arun_{session_num}folder for each trial to keep logs isolated.Update log path handling in
train_test_model
Pass the unique run directory to your model training function, so TensorBoard writes logs to that specific folder.Simplify histogram naming (optional but recommended)
Instead of using sequential numbers (W1, B1), use the actual layer names from your model—this makes it easier to map histograms to your model structure.
Modified Full Code
# SET UP import tensorflow as tf from tensorflow import keras from tensorboard.plugins.hparams import api as hp import os keras.backend.clear_session() # LOAD DATA (X_train_full, y_train_full), (X_test, y_test) = keras.datasets.mnist.load_data() X_train_reduced = X_train_full[y_train_full < 5] y_train_reduced = y_train_full[y_train_full < 5] X_test_reduced = X_test[y_test < 5] y_test_reduced = y_test[y_test < 5] X_train = X_train_reduced[5000:] y_train = y_train_reduced[5000:] X_valid = X_train_reduced[:5000] y_valid = y_train_reduced[:5000] # Set hyperparameters to tune keras.backend.clear_session() HP_INITIAL_LEARNING_RATE = hp.HParam("initial_learning_rate", hp.Discrete([0.0001, 0.00012, 0.00015])) HP_NUM_BATCH_SIZE = hp.HParam("batch_size", hp.Discrete([32])) HP_NUM_EPOCHS = hp.HParam("epochs", hp.Discrete([180])) HP_BETA_1 = hp.HParam("beta_1", hp.Discrete([0.95])) HP_BETA_2 = hp.HParam("beta_2", hp.Discrete([0.9994])) HP_DECAY_STEP = hp.HParam("decay_step", hp.Discrete([10000])) HP_DECAY_RATE = hp.HParam("decay_rate", hp.Discrete([0.8])) # Function to create summaries def run(run_dir, hparams): with tf.summary.create_file_writer(run_dir).as_default() as summ: hp.hparams(hparams) # Record hyperparameters for this trial num_layers = 5 model = train_test_model(hparams, run_dir, num_layers=num_layers) # Log histograms using layer names instead of sequential numbers for layer in model.layers[1:]: layer_name = layer.name if isinstance(layer, keras.layers.BatchNormalization): weights, biases, _, _ = layer.get_weights() tf.summary.histogram(name=f"{layer_name}/weights", data=weights, step=1, description=f"Weights of {layer_name}") tf.summary.histogram(name=f"{layer_name}/biases", data=biases, step=1, description=f"Biases of {layer_name}") elif isinstance(layer, keras.layers.Dense): weights, biases = layer.get_weights() tf.summary.histogram(name=f"{layer_name}/weights", data=weights, step=1, description=f"Weights of {layer_name}") tf.summary.histogram(name=f"{layer_name}/biases", data=biases, step=1, description=f"Biases of {layer_name}") # Function to train/test model with given hyperparameters def train_test_model(hparams, run_dir, num_layers=5): act_fun = "elu" initializer = "he_normal" # Create a descriptive name for the model run run_name = f"MNIST_0-4-nDens_{num_layers}-{act_fun}-{initializer}-btch_sz_{hparams[HP_NUM_BATCH_SIZE]}-epo_{hparams[HP_NUM_EPOCHS]}-Adam-b1_{hparams[HP_BETA_1]}-b2_{hparams[HP_BETA_2]}-lr_{hparams[HP_INITIAL_LEARNING_RATE]}-decay_{hparams[HP_DECAY_RATE]}-step_{hparams[HP_DECAY_STEP]}" # Full log directory for this specific run logdir = os.path.join(run_dir, run_name) os.makedirs(logdir, exist_ok=True) # Create folder if it doesn't exist # Build the model model = keras.models.Sequential() model.add(keras.layers.Flatten(input_shape=[28, 28])) model.add(keras.layers.BatchNormalization()) hidden_layer_neurons = 100 for _ in range(num_layers): model.add(keras.layers.Dense(hidden_layer_neurons, activation=act_fun, kernel_initializer="he_normal")) model.add(keras.layers.BatchNormalization()) model.add(keras.layers.Dense(5, activation='softmax')) # Set up optimizer and callbacks lr_schedule = keras.optimizers.schedules.ExponentialDecay( initial_learning_rate=hparams[HP_INITIAL_LEARNING_RATE], decay_steps=hparams[HP_DECAY_STEP], decay_rate=hparams[HP_DECAY_RATE]) optimizer = keras.optimizers.Adam(learning_rate=lr_schedule, beta_1=hparams[HP_BETA_1], beta_2=hparams[HP_BETA_2]) early_stopping_cb = keras.callbacks.EarlyStopping(patience=10, restore_best_weights=True) tensorboard_cb = keras.callbacks.TensorBoard(logdir, histogram_freq=1) callbacks = [ early_stopping_cb, tensorboard_cb, hp.KerasCallback(logdir, hparams), # Link HPARAMs to this run's logs ] # Compile and train model.compile(loss='sparse_categorical_crossentropy', metrics=["accuracy"], optimizer=optimizer) model.fit(X_train, y_train, epochs=hparams[HP_NUM_EPOCHS], batch_size=hparams[HP_NUM_BATCH_SIZE], validation_data=(X_valid, y_valid), callbacks=callbacks) return model # Run all hyperparameter combinations session_num = 0 ROOT_LOG_FOLDER = 'logs/batch_normalization/' for batch_size in HP_NUM_BATCH_SIZE.domain.values: for epochs in HP_NUM_EPOCHS.domain.values: for beta_1 in HP_BETA_1.domain.values: for beta_2 in HP_BETA_2.domain.values: for lr in HP_INITIAL_LEARNING_RATE.domain.values: for decay_rate in HP_DECAY_RATE.domain.values: for decay_step in HP_DECAY_STEP.domain.values: hparams = { HP_NUM_BATCH_SIZE: batch_size, HP_NUM_EPOCHS: epochs, HP_BETA_1: beta_1, HP_BETA_2: beta_2, HP_INITIAL_LEARNING_RATE: lr, HP_DECAY_RATE: decay_rate, HP_DECAY_STEP: decay_step } # Create unique run directory current_run_dir = os.path.join(ROOT_LOG_FOLDER, f"run_{session_num}") print(f"Starting trial {session_num}: { {h.name: hparams[h] for h in hparams} }") run(current_run_dir, hparams) session_num += 1
Key Changes Explained
- Isolated Log Directories: Each hyperparameter trial gets its own
run_{session_num}folder. TensorBoard automatically recognizes these as separate runs. - Layer-Based Histogram Names: Using
layer.nameinstead of sequential numbers makes it easy to match histograms to specific layers in your model. - HPARAM Integration: The
hp.KerasCallbacklinks each run's hyperparameters to its logs, so you can filter runs by hyperparameters in the TensorBoard HPARAMs tab and jump directly to their histograms.
What to Expect in TensorBoard
After running the modified code:
- Open the HPARAMS tab to see all your trials, with hyperparameters and metrics listed. Click any trial to view its details.
- In the HISTOGRAMS or DISTRIBUTIONS tab, use the "Runs" dropdown in the top-right to switch between different model runs. Each run's weight/bias histograms will be grouped separately, no more merged sequences!
内容的提问来源于stack exchange,提问作者roschach

