能否在Keras中构建输出反缩放自定义层?Keras回归网络缩放逻辑内置及输出训练策略咨询
Great question! Handling feature scaling in a way that’s integrated directly into your model is totally feasible, and it’ll eliminate the need to manually save/load means and max absolute values for inference. Plus, we’ll address your question about output scaling to make sure your model trains effectively across all variable magnitudes.
1. Building Custom Layers for Scaling and Inverse Scaling
The key idea is to create two custom Keras layers: one to scale your input data using your max-mean formula, and another to inverse-scale your model’s predictions back to the original magnitude. These layers will store the mean and max absolute values from your training data as non-trainable variables, so they’ll be saved alongside your model—no separate files needed.
Step 1: Define the Custom Scaling Layers
import tensorflow as tf from tensorflow.keras.layers import Layer import numpy as np class InputMaxMeanScaler(Layer): def __init__(self, means, max_abss, **kwargs): super().__init__(**kwargs) # Store training stats as non-trainable variables (saved with the model) self.means = tf.Variable(means, dtype=tf.float32, trainable=False) self.max_abss = tf.Variable(max_abss, dtype=tf.float32, trainable=False) def call(self, inputs): # Apply your scaling formula: (x - mean) / max(|X|) return (inputs - self.means) / self.max_abss class OutputMaxMeanInverseScaler(Layer): def __init__(self, means, max_abss, **kwargs): super().__init__(**kwargs) self.means = tf.Variable(means, dtype=tf.float32, trainable=False) self.max_abss = tf.Variable(max_abss, dtype=tf.float32, trainable=False) def call(self, inputs): # Reverse scaling: x' * max(|X|) + mean return inputs * self.max_abss + self.means
Step 2: Calculate Training Data Statistics
First, compute the mean and max absolute value for each input and output feature from your training dataset. Make sure to handle cases where max_abss could be 0 (e.g., a feature with identical values across all samples) to avoid division by zero.
# Assume X_train is your input data (shape: [num_samples, 60]) # y_train is your target data (shape: [num_samples, 35]) # Compute input stats mean_in = np.mean(X_train, axis=0) max_abs_in = np.max(np.abs(X_train), axis=0) # Replace 0s with a small epsilon to prevent division by zero max_abs_in = np.where(max_abs_in == 0, 1e-8, max_abs_in) # Compute output stats mean_out = np.mean(y_train, axis=0) max_abs_out = np.max(np.abs(y_train), axis=0) max_abs_out = np.where(max_abs_out == 0, 1e-8, max_abs_out)
Step 3: Build and Train Your Model
We’ll create two versions of the model: one for training (outputs scaled values) and one for inference (outputs original-magnitude values).
Training Model
This model takes raw input, scales it, and outputs scaled predictions (matching your scaled target values):
# Input layer input_layer = tf.keras.Input(shape=(60,)) # Apply input scaling scaled_input = InputMaxMeanScaler(mean_in, max_abs_in)(input_layer) # Your core regression network (customize this to your needs) x = tf.keras.layers.Dense(128, activation='relu')(scaled_input) x = tf.keras.layers.Dense(64, activation='relu')(x) # Output scaled values (in [-1, 1] range) scaled_output = tf.keras.layers.Dense(35)(x) # Compile and train train_model = tf.keras.Model(input_layer, scaled_output) train_model.compile(optimizer='adam', loss='mse') # Scale your target values before training scaled_y_train = (y_train - mean_out) / max_abs_out train_model.fit(X_train, scaled_y_train, epochs=50, batch_size=32, validation_split=0.1)
Inference Model
Now extend the trained model with the inverse scaling layer to get predictions in the original magnitude. This is the model you’ll save and use for deployment:
# Add inverse scaling to the trained network raw_predictions = OutputMaxMeanInverseScaler(mean_out, max_abs_out)(scaled_output) inference_model = tf.keras.Model(input_layer, raw_predictions) # Save the model (includes scaling stats!) inference_model.save('scaled_regression_model.h5') # Later, load and use it without needing separate stats files loaded_model = tf.keras.models.load_model( 'scaled_regression_model.h5', custom_objects={ 'InputMaxMeanScaler': InputMaxMeanScaler, 'OutputMaxMeanInverseScaler': OutputMaxMeanInverseScaler } ) # Predict directly on raw input data test_predictions = loaded_model.predict(X_test)
2. Should You Skip Output Scaling?
No—you absolutely should scale your output variables. Here’s why:
- Magnitude bias: Without scaling, your loss function will be dominated by errors from the 100k-scale variables. For example, an error of 100 on a 100k variable contributes 10,000 to the MSE loss, while an error of 1 on a ±10 variable contributes just 1. Your model will prioritize fitting the large-scale variables and ignore the small ones entirely.
- Training stability: Scaling outputs to [-1, 1] keeps gradient magnitudes consistent across all outputs, preventing issues like exploding gradients and making training converge faster and more reliably.
Skipping output scaling will almost certainly lead to poor performance on your smaller-magnitude target variables, so don’t skip this step!
内容的提问来源于stack exchange,提问作者Jenna

