You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

tf.keras线性回归调优疑问:RMSE计算、迭代参数与过程解析

Linear Regression Training: Iteration Logic, Loss Mismatches, and Parameter Tracking

Hey there! Let's work through your questions one by one since you're diving into linear regression with synthetic data and hitting some confusing spots with training loops, loss calculations, and parameter tracking.

1. Understanding Epochs, Batches, and Iterations

First, let's clarify the core definitions to untangle their relationship:

  • Epoch: One full pass through your entire training dataset. If you have 12 samples, one epoch means the model sees all 12 samples once.
  • Batch Size: The number of samples the model processes in a single gradient update step. For your setup with batch_size=4, each step uses 4 samples.
  • Iteration: A single gradient update step (processing one batch). For 12 samples and batch size 4, each epoch has 3 iterations (12 / 4 = 3). With epochs=4, you get 4 * 3 = 12 total iterations—each iteration tweaks the model's parameters once based on the batch's data.

So your 12 total iterations make perfect sense: 4 full passes over the data, each split into 3 batches.

2. Why Your Manual RMSE/Loss Doesn't Match tf.keras?

The mismatch between your manual calculations (RMSE=21.8, Loss=476.82) and tf.keras outputs (RMSE=23.1, Loss=535.48) usually boils down to a few key differences:

Common Culprits:

  • Model Parameters vs. Closed-Form Solution: If you calculated using the linear regression closed-form (least squares) solution, that's the optimal fit for the entire dataset. But tf.keras uses stochastic gradient descent (SGD) with batches, so the final parameters are the result of incremental updates over batches/epochs—they might not match the closed-form solution yet (especially if epochs are low or learning rate is suboptimal).
  • Loss Calculation Details:
    • tf.keras's MeanSquaredError loss takes the average of squared errors across the batch (or entire dataset for evaluation), not the sum. Double-check if you divided by the number of samples when calculating your manual loss.
    • Floating-point precision: tf.keras uses float32 by default, while Excel uses float64. Small rounding differences can add up over calculations.
  • Data Preprocessing: Did the Colab notebook normalize/standardize the input data? If the model was trained on scaled X values but you used raw X for manual calculations, your predictions will be way off.

How to Debug:

  1. Grab the model's final parameters with:
    slope, intercept = model.layers[0].get_weights()
    slope = slope[0][0]
    intercept = intercept[0]
    
  2. Use these exact values to recalculate your manual RMSE/loss:
    • For each sample, compute y_pred = slope * X + intercept
    • Calculate squared error for each sample: (y_pred - y_true)²
    • Compute MSE as the average of these squared errors, then RMSE as the square root of MSE.
  3. Compare this to tf.keras's output using:
    mse, rmse = model.evaluate(X, y, verbose=0)
    print(f"tf.keras MSE: {mse}, RMSE: {rmse}")
    

This should align if you're using the same parameters and data.

3. Tracking Parameters and RMSE Every Iteration

To capture slope, intercept, and RMSE after every batch iteration (not just per epoch), you'll need to use a custom Keras Callback. Here's a step-by-step implementation:

Step 1: Define the Custom Callback

This callback will log parameters and metrics after each batch:

import tensorflow as tf
import numpy as np

class BatchTracker(tf.keras.callbacks.Callback):
    def __init__(self,Fill X Rub.Farin                file(费):dr,苦ermaintenance_jump to the code,哦不对,重新写:
    def __init__(self, X, y):
        super().__init__()
        self.X = X  # Full dataset (to compute full RMSE if needed)
        self.y = y
        self.slopes = []
        self.intercepts = []
        self.batch_rmse = []
        self.full_dataset_rmse = []

    def on_train_batch_end(self, batch, logs=None):
        # Grab current model weights
        w, b = self.model.layers[0].get_weights()
        self.slopes.append(w[0][0])
        self.intercepts.append(b[0])
        
        # Log RMSE from the current batch (provided by logs)
        self.batch_rmse.append(logs['root_mean_squared_error'])
        
        # Optional: Compute RMSE on the full dataset after this iteration
        y_pred = self.model.predict(self.X, verbose=0)
        full_rmse = np.sqrt(np.mean((y_pred - self.y)**2))
        self.full_dataset_rmse.append(full_rmse)

Step 2: Use the Callback During Training

# Assume your model and data are already defined
tracker = BatchTracker(X, y)

# Train with the callback
history = model.fit(
    X, y,
    epochs=4,
    batch_size=4,
    callbacks=[tracker],
    verbose=1
)

# Access the tracked data
print("All slopes per iteration:", tracker.slopes)
print("All intercepts per iteration:", tracker.intercepts)
print("RMSE per batch:", tracker.batch_rmse)
print("Full dataset RMSE per iteration:", tracker.full_dataset_rmse)

Bonus: Manual Gradient Descent for Full Control

If you want to fully trace every parameter update (including initial random values), implement small-batch SGD manually with tf.GradientTape:

# Initialize random parameters (matches Keras's default initialization)
w = tf.Variable(np.random.randn(1), dtype=tf.float32)
b = tf.Variable(np.random.randn(1), dtype=tf.float32)
learning_rate = 0.01
epochs = 4
batch_size = 4
num_samples = len(X)
num_batches = num_samples // batch_size

# Track updates
slopes = [w.numpy()[0]]
intercepts = [b.numpy()[0]]
batch_losses = []

for epoch in range(epochs):
    # Shuffle data (mimics SGD's randomness)
    shuffled_indices = np.random.permutation(num_samples)
    X_shuffled = X[shuffled_indices]
    y_shuffled = y[shuffled_indices]
    
    for batch in range(num_batches):
        start = batch * batch_size
        end = start + batch_size
        X_batch = X_shuffled[start:end]
        y_batch = y_shuffled[start:end]
        
        # Compute loss and gradients
        with tf.GradientTape() as tape:
            y_pred = tf.multiply(X_batch, w) + b
            loss = tf.reduce_mean(tf.square(y_pred - y_batch))
        
        # Calculate gradients
        dw, db = tape.gradient(loss, [w, b])
        
        # Update parameters
        w.assign_sub(learning_rate * dw)
        b.assign_sub(learning_rate * db)
        
        # Log values
        slopes.append(w.numpy()[0])
        intercepts.append(b.numpy()[0])
        batch_losses.append(loss.numpy())

print("Initial slope:", slopes[0])
print("Initial intercept:", intercepts[0])
print("All parameter updates (slope, intercept):", list(zip(slopes, intercepts)))

This gives you full visibility into how each batch's gradient adjusts the parameters, exactly as Keras does under the hood.

内容的提问来源于stack exchange,提问作者Stat Tistician

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 07:07:47