tf.keras线性回归调优疑问:RMSE计算、迭代参数与过程解析
Hey there! Let's work through your questions one by one since you're diving into linear regression with synthetic data and hitting some confusing spots with training loops, loss calculations, and parameter tracking.
1. Understanding Epochs, Batches, and Iterations
First, let's clarify the core definitions to untangle their relationship:
- Epoch: One full pass through your entire training dataset. If you have 12 samples, one epoch means the model sees all 12 samples once.
- Batch Size: The number of samples the model processes in a single gradient update step. For your setup with
batch_size=4, each step uses 4 samples. - Iteration: A single gradient update step (processing one batch). For 12 samples and batch size 4, each epoch has 3 iterations (12 / 4 = 3). With
epochs=4, you get 4 * 3 = 12 total iterations—each iteration tweaks the model's parameters once based on the batch's data.
So your 12 total iterations make perfect sense: 4 full passes over the data, each split into 3 batches.
2. Why Your Manual RMSE/Loss Doesn't Match tf.keras?
The mismatch between your manual calculations (RMSE=21.8, Loss=476.82) and tf.keras outputs (RMSE=23.1, Loss=535.48) usually boils down to a few key differences:
Common Culprits:
- Model Parameters vs. Closed-Form Solution: If you calculated using the linear regression closed-form (least squares) solution, that's the optimal fit for the entire dataset. But tf.keras uses stochastic gradient descent (SGD) with batches, so the final parameters are the result of incremental updates over batches/epochs—they might not match the closed-form solution yet (especially if epochs are low or learning rate is suboptimal).
- Loss Calculation Details:
- tf.keras's
MeanSquaredErrorloss takes the average of squared errors across the batch (or entire dataset for evaluation), not the sum. Double-check if you divided by the number of samples when calculating your manual loss. - Floating-point precision: tf.keras uses float32 by default, while Excel uses float64. Small rounding differences can add up over calculations.
- tf.keras's
- Data Preprocessing: Did the Colab notebook normalize/standardize the input data? If the model was trained on scaled X values but you used raw X for manual calculations, your predictions will be way off.
How to Debug:
- Grab the model's final parameters with:
slope, intercept = model.layers[0].get_weights() slope = slope[0][0] intercept = intercept[0] - Use these exact values to recalculate your manual RMSE/loss:
- For each sample, compute
y_pred = slope * X + intercept - Calculate squared error for each sample:
(y_pred - y_true)² - Compute MSE as the average of these squared errors, then RMSE as the square root of MSE.
- For each sample, compute
- Compare this to tf.keras's output using:
mse, rmse = model.evaluate(X, y, verbose=0) print(f"tf.keras MSE: {mse}, RMSE: {rmse}")
This should align if you're using the same parameters and data.
3. Tracking Parameters and RMSE Every Iteration
To capture slope, intercept, and RMSE after every batch iteration (not just per epoch), you'll need to use a custom Keras Callback. Here's a step-by-step implementation:
Step 1: Define the Custom Callback
This callback will log parameters and metrics after each batch:
import tensorflow as tf import numpy as np class BatchTracker(tf.keras.callbacks.Callback): def __init__(self,Fill X Rub.Farin file(费):dr,苦ermaintenance_jump to the code,哦不对,重新写: def __init__(self, X, y): super().__init__() self.X = X # Full dataset (to compute full RMSE if needed) self.y = y self.slopes = [] self.intercepts = [] self.batch_rmse = [] self.full_dataset_rmse = [] def on_train_batch_end(self, batch, logs=None): # Grab current model weights w, b = self.model.layers[0].get_weights() self.slopes.append(w[0][0]) self.intercepts.append(b[0]) # Log RMSE from the current batch (provided by logs) self.batch_rmse.append(logs['root_mean_squared_error']) # Optional: Compute RMSE on the full dataset after this iteration y_pred = self.model.predict(self.X, verbose=0) full_rmse = np.sqrt(np.mean((y_pred - self.y)**2)) self.full_dataset_rmse.append(full_rmse)
Step 2: Use the Callback During Training
# Assume your model and data are already defined tracker = BatchTracker(X, y) # Train with the callback history = model.fit( X, y, epochs=4, batch_size=4, callbacks=[tracker], verbose=1 ) # Access the tracked data print("All slopes per iteration:", tracker.slopes) print("All intercepts per iteration:", tracker.intercepts) print("RMSE per batch:", tracker.batch_rmse) print("Full dataset RMSE per iteration:", tracker.full_dataset_rmse)
Bonus: Manual Gradient Descent for Full Control
If you want to fully trace every parameter update (including initial random values), implement small-batch SGD manually with tf.GradientTape:
# Initialize random parameters (matches Keras's default initialization) w = tf.Variable(np.random.randn(1), dtype=tf.float32) b = tf.Variable(np.random.randn(1), dtype=tf.float32) learning_rate = 0.01 epochs = 4 batch_size = 4 num_samples = len(X) num_batches = num_samples // batch_size # Track updates slopes = [w.numpy()[0]] intercepts = [b.numpy()[0]] batch_losses = [] for epoch in range(epochs): # Shuffle data (mimics SGD's randomness) shuffled_indices = np.random.permutation(num_samples) X_shuffled = X[shuffled_indices] y_shuffled = y[shuffled_indices] for batch in range(num_batches): start = batch * batch_size end = start + batch_size X_batch = X_shuffled[start:end] y_batch = y_shuffled[start:end] # Compute loss and gradients with tf.GradientTape() as tape: y_pred = tf.multiply(X_batch, w) + b loss = tf.reduce_mean(tf.square(y_pred - y_batch)) # Calculate gradients dw, db = tape.gradient(loss, [w, b]) # Update parameters w.assign_sub(learning_rate * dw) b.assign_sub(learning_rate * db) # Log values slopes.append(w.numpy()[0]) intercepts.append(b.numpy()[0]) batch_losses.append(loss.numpy()) print("Initial slope:", slopes[0]) print("Initial intercept:", intercepts[0]) print("All parameter updates (slope, intercept):", list(zip(slopes, intercepts)))
This gives you full visibility into how each batch's gradient adjusts the parameters, exactly as Keras does under the hood.
内容的提问来源于stack exchange,提问作者Stat Tistician

