You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Keras的TensorBoard中按模型分组直方图?

Fixing TensorBoard Histogram Grouping for HPARAM-Tuned Keras Models

Hey there, let's tackle this histogram grouping issue you're hitting with TensorBoard during your MNIST classification HPARAM tuning. The core problem here is that you're writing all your model logs into the same root directory—TensorBoard merges all data under the same folder, which is why your weight/bias histograms are showing up as a single continuous sequence (W1-W6, W7-W12, etc.) instead of being grouped per model.

The Right Approach: Independent Log Directories per Run

TensorBoard relies on separate subdirectories to distinguish different training runs. This is also the standard way to work with the HPARAM plugin, since it ties each set of hyperparameters to its own log folder. Adding tf.name_scope won't fix this because it only adds a prefix within the same log context—you need full directory separation.

Step-by-Step Code Modifications

Here's how to adjust your code to get proper grouping:

  1. Generate a unique log subdirectory for each hyperparameter trial
    We'll create a run_{session_num} folder for each trial to keep logs isolated.

  2. Update log path handling in train_test_model
    Pass the unique run directory to your model training function, so TensorBoard writes logs to that specific folder.

  3. Simplify histogram naming (optional but recommended)
    Instead of using sequential numbers (W1, B1), use the actual layer names from your model—this makes it easier to map histograms to your model structure.

Modified Full Code

# SET UP
import tensorflow as tf
from tensorflow import keras
from tensorboard.plugins.hparams import api as hp
import os
keras.backend.clear_session()

# LOAD DATA
(X_train_full, y_train_full), (X_test, y_test) = keras.datasets.mnist.load_data()
X_train_reduced = X_train_full[y_train_full < 5]
y_train_reduced = y_train_full[y_train_full < 5]
X_test_reduced = X_test[y_test < 5]
y_test_reduced = y_test[y_test < 5]
X_train = X_train_reduced[5000:]
y_train = y_train_reduced[5000:]
X_valid = X_train_reduced[:5000]
y_valid = y_train_reduced[:5000]

# Set hyperparameters to tune
keras.backend.clear_session()
HP_INITIAL_LEARNING_RATE = hp.HParam("initial_learning_rate", hp.Discrete([0.0001, 0.00012, 0.00015]))
HP_NUM_BATCH_SIZE = hp.HParam("batch_size", hp.Discrete([32]))
HP_NUM_EPOCHS = hp.HParam("epochs", hp.Discrete([180]))
HP_BETA_1 = hp.HParam("beta_1", hp.Discrete([0.95]))
HP_BETA_2 = hp.HParam("beta_2", hp.Discrete([0.9994]))
HP_DECAY_STEP = hp.HParam("decay_step", hp.Discrete([10000]))
HP_DECAY_RATE = hp.HParam("decay_rate", hp.Discrete([0.8]))

# Function to create summaries
def run(run_dir, hparams):
    with tf.summary.create_file_writer(run_dir).as_default() as summ:
        hp.hparams(hparams)  # Record hyperparameters for this trial
        num_layers = 5
        model = train_test_model(hparams, run_dir, num_layers=num_layers)
        
        # Log histograms using layer names instead of sequential numbers
        for layer in model.layers[1:]:
            layer_name = layer.name
            if isinstance(layer, keras.layers.BatchNormalization):
                weights, biases, _, _ = layer.get_weights()
                tf.summary.histogram(name=f"{layer_name}/weights", data=weights, step=1, description=f"Weights of {layer_name}")
                tf.summary.histogram(name=f"{layer_name}/biases", data=biases, step=1, description=f"Biases of {layer_name}")
            elif isinstance(layer, keras.layers.Dense):
                weights, biases = layer.get_weights()
                tf.summary.histogram(name=f"{layer_name}/weights", data=weights, step=1, description=f"Weights of {layer_name}")
                tf.summary.histogram(name=f"{layer_name}/biases", data=biases, step=1, description=f"Biases of {layer_name}")

# Function to train/test model with given hyperparameters
def train_test_model(hparams, run_dir, num_layers=5):
    act_fun = "elu"
    initializer = "he_normal"
    # Create a descriptive name for the model run
    run_name = f"MNIST_0-4-nDens_{num_layers}-{act_fun}-{initializer}-btch_sz_{hparams[HP_NUM_BATCH_SIZE]}-epo_{hparams[HP_NUM_EPOCHS]}-Adam-b1_{hparams[HP_BETA_1]}-b2_{hparams[HP_BETA_2]}-lr_{hparams[HP_INITIAL_LEARNING_RATE]}-decay_{hparams[HP_DECAY_RATE]}-step_{hparams[HP_DECAY_STEP]}"
    # Full log directory for this specific run
    logdir = os.path.join(run_dir, run_name)
    os.makedirs(logdir, exist_ok=True)  # Create folder if it doesn't exist

    # Build the model
    model = keras.models.Sequential()
    model.add(keras.layers.Flatten(input_shape=[28, 28]))
    model.add(keras.layers.BatchNormalization())
    hidden_layer_neurons = 100
    for _ in range(num_layers):
        model.add(keras.layers.Dense(hidden_layer_neurons, activation=act_fun, kernel_initializer="he_normal"))
        model.add(keras.layers.BatchNormalization())
    model.add(keras.layers.Dense(5, activation='softmax'))

    # Set up optimizer and callbacks
    lr_schedule = keras.optimizers.schedules.ExponentialDecay(
        initial_learning_rate=hparams[HP_INITIAL_LEARNING_RATE],
        decay_steps=hparams[HP_DECAY_STEP],
        decay_rate=hparams[HP_DECAY_RATE])
    optimizer = keras.optimizers.Adam(learning_rate=lr_schedule, beta_1=hparams[HP_BETA_1], beta_2=hparams[HP_BETA_2])

    early_stopping_cb = keras.callbacks.EarlyStopping(patience=10, restore_best_weights=True)
    tensorboard_cb = keras.callbacks.TensorBoard(logdir, histogram_freq=1)
    callbacks = [
        early_stopping_cb,
        tensorboard_cb,
        hp.KerasCallback(logdir, hparams),  # Link HPARAMs to this run's logs
    ]

    # Compile and train
    model.compile(loss='sparse_categorical_crossentropy', metrics=["accuracy"], optimizer=optimizer)
    model.fit(X_train, y_train, epochs=hparams[HP_NUM_EPOCHS], batch_size=hparams[HP_NUM_BATCH_SIZE],
              validation_data=(X_valid, y_valid), callbacks=callbacks)
    return model

# Run all hyperparameter combinations
session_num = 0
ROOT_LOG_FOLDER = 'logs/batch_normalization/'
for batch_size in HP_NUM_BATCH_SIZE.domain.values:
    for epochs in HP_NUM_EPOCHS.domain.values:
        for beta_1 in HP_BETA_1.domain.values:
            for beta_2 in HP_BETA_2.domain.values:
                for lr in HP_INITIAL_LEARNING_RATE.domain.values:
                    for decay_rate in HP_DECAY_RATE.domain.values:
                        for decay_step in HP_DECAY_STEP.domain.values:
                            hparams = {
                                HP_NUM_BATCH_SIZE: batch_size,
                                HP_NUM_EPOCHS: epochs,
                                HP_BETA_1: beta_1,
                                HP_BETA_2: beta_2,
                                HP_INITIAL_LEARNING_RATE: lr,
                                HP_DECAY_RATE: decay_rate,
                                HP_DECAY_STEP: decay_step
                            }
                            # Create unique run directory
                            current_run_dir = os.path.join(ROOT_LOG_FOLDER, f"run_{session_num}")
                            print(f"Starting trial {session_num}: { {h.name: hparams[h] for h in hparams} }")
                            run(current_run_dir, hparams)
                            session_num += 1

Key Changes Explained

  • Isolated Log Directories: Each hyperparameter trial gets its own run_{session_num} folder. TensorBoard automatically recognizes these as separate runs.
  • Layer-Based Histogram Names: Using layer.name instead of sequential numbers makes it easy to match histograms to specific layers in your model.
  • HPARAM Integration: The hp.KerasCallback links each run's hyperparameters to its logs, so you can filter runs by hyperparameters in the TensorBoard HPARAMs tab and jump directly to their histograms.

What to Expect in TensorBoard

After running the modified code:

  1. Open the HPARAMS tab to see all your trials, with hyperparameters and metrics listed. Click any trial to view its details.
  2. In the HISTOGRAMS or DISTRIBUTIONS tab, use the "Runs" dropdown in the top-right to switch between different model runs. Each run's weight/bias histograms will be grouped separately, no more merged sequences!

内容的提问来源于stack exchange,提问作者roschach

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 08:47:43