简单TensorFlow示例中网络发散并出现NaN问题求助
Hey there, I’ve dealt with this exact NaN/divergence issue when training regression models before—total headache! Let’s walk through the most common fixes for your quadratic regression (Y = WX² + uX + b) using the slr05.xls dataset:
Common Causes & Solutions
1. Normalize/Standardize Your Features (Critical!)
The quadratic term X² can easily blow up to huge values, which makes gradient updates spiral out of control. Normalizing all features (including X²) to a similar numerical range is usually the first fix to try.
Here’s how to implement it with scikit-learn (or you can do it manually):
import numpy as np from sklearn.preprocessing import StandardScaler # Extract raw features and labels from your dataset X_raw = dataset[:, 0].reshape(-1, 1) y_raw = dataset[:, 1].reshape(-1, 1) # Create quadratic feature X_squared = X_raw ** 2 # Combine quadratic and linear features features = np.hstack([X_squared, X_raw]) # Standardize features (mean=0, std=1) scaler_X = StandardScaler() scaled_features = scaler_X.fit_transform(features) # Optional: Standardize labels too to keep loss values manageable scaler_y = StandardScaler() scaled_y = scaler_y.fit_transform(y_raw)
After training, don’t forget to inverse-transform your predictions to get back to the original scale:
# Example: After getting predicted values from your model predicted_y = scaler_y.inverse_transform(model_predictions)
2. Lower Your Learning Rate
If your optimizer’s learning rate is too high, the weight updates will overshoot the optimal values, leading to divergence and NaNs. Try scaling it down drastically—start with something like 0.001 instead of 0.1 or 0.01.
You can also switch to an adaptive optimizer like Adam, which automatically adjusts learning rates during training:
# Replace SGD with Adam (usually more stable) optimizer = tf.optimizers.Adam(learning_rate=0.001)
3. Initialize Weights Properly
Avoid starting with large random weights—they can cause huge initial loss values and unstable gradients. Use small, constrained initializations:
# Small random values with low standard deviation W = tf.Variable(tf.random.normal(shape=[1], mean=0.0, stddev=0.01)) u = tf.Variable(tf.random.normal(shape=[1], mean=0.0, stddev=0.01)) b = tf.Variable(tf.zeros(shape=[1])) # Bias starts at 0 is safe
Or use Glorot/Xavier initialization, which is designed to keep gradients stable across layers:
initializer = tf.initializers.GlorotUniform() W = tf.Variable(initializer(shape=[1])) u = tf.Variable(initializer(shape=[1]))
4. Add Gradient Clipping
If gradients get too large, clip them to a maximum norm to prevent explosion:
@tf.function def train_step(x, y): with tf.GradientTape() as tape: # Assume x is your scaled [X², X] features y_pred = W * x[:, 0:1] + u * x[:, 1:2] + b loss = tf.reduce_mean(tf.square(y_pred - y)) gradients = tape.gradient(loss, [W, u, b]) # Clip gradients to max norm of 5.0 (adjust this value if needed) clipped_gradients = [tf.clip_by_norm(grad, 5.0) for grad in gradients] optimizer.apply_gradients(zip(clipped_gradients, [W, u, b])) return loss
5. Check for Outliers in Your Dataset
Open up slr05.xls and scan for extreme values in X or Y—outliers can cause massive loss spikes that break training. You can either remove these outliers or replace them with median values to stabilize training.
Quick Troubleshooting Order
Start with feature normalization → then lower learning rate/switch to Adam → if still broken, try gradient clipping or weight initialization tweaks. Most of the time, normalization alone fixes the NaN issue!
内容的提问来源于stack exchange,提问作者Jespar

