You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

简单TensorFlow示例中网络发散并出现NaN问题求助

Fixing NaN/Divergence in TensorFlow Quadratic Regression Model

Hey there, I’ve dealt with this exact NaN/divergence issue when training regression models before—total headache! Let’s walk through the most common fixes for your quadratic regression (Y = WX² + uX + b) using the slr05.xls dataset:

Common Causes & Solutions

1. Normalize/Standardize Your Features (Critical!)

The quadratic term X² can easily blow up to huge values, which makes gradient updates spiral out of control. Normalizing all features (including X²) to a similar numerical range is usually the first fix to try.

Here’s how to implement it with scikit-learn (or you can do it manually):

import numpy as np
from sklearn.preprocessing import StandardScaler

# Extract raw features and labels from your dataset
X_raw = dataset[:, 0].reshape(-1, 1)
y_raw = dataset[:, 1].reshape(-1, 1)

# Create quadratic feature
X_squared = X_raw ** 2
# Combine quadratic and linear features
features = np.hstack([X_squared, X_raw])

# Standardize features (mean=0, std=1)
scaler_X = StandardScaler()
scaled_features = scaler_X.fit_transform(features)

# Optional: Standardize labels too to keep loss values manageable
scaler_y = StandardScaler()
scaled_y = scaler_y.fit_transform(y_raw)

After training, don’t forget to inverse-transform your predictions to get back to the original scale:

# Example: After getting predicted values from your model
predicted_y = scaler_y.inverse_transform(model_predictions)

2. Lower Your Learning Rate

If your optimizer’s learning rate is too high, the weight updates will overshoot the optimal values, leading to divergence and NaNs. Try scaling it down drastically—start with something like 0.001 instead of 0.1 or 0.01.

You can also switch to an adaptive optimizer like Adam, which automatically adjusts learning rates during training:

# Replace SGD with Adam (usually more stable)
optimizer = tf.optimizers.Adam(learning_rate=0.001)

3. Initialize Weights Properly

Avoid starting with large random weights—they can cause huge initial loss values and unstable gradients. Use small, constrained initializations:

# Small random values with low standard deviation
W = tf.Variable(tf.random.normal(shape=[1], mean=0.0, stddev=0.01))
u = tf.Variable(tf.random.normal(shape=[1], mean=0.0, stddev=0.01))
b = tf.Variable(tf.zeros(shape=[1]))  # Bias starts at 0 is safe

Or use Glorot/Xavier initialization, which is designed to keep gradients stable across layers:

initializer = tf.initializers.GlorotUniform()
W = tf.Variable(initializer(shape=[1]))
u = tf.Variable(initializer(shape=[1]))

4. Add Gradient Clipping

If gradients get too large, clip them to a maximum norm to prevent explosion:

@tf.function
def train_step(x, y):
    with tf.GradientTape() as tape:
        # Assume x is your scaled [X², X] features
        y_pred = W * x[:, 0:1] + u * x[:, 1:2] + b
        loss = tf.reduce_mean(tf.square(y_pred - y))
    
    gradients = tape.gradient(loss, [W, u, b])
    # Clip gradients to max norm of 5.0 (adjust this value if needed)
    clipped_gradients = [tf.clip_by_norm(grad, 5.0) for grad in gradients]
    optimizer.apply_gradients(zip(clipped_gradients, [W, u, b]))
    
    return loss

5. Check for Outliers in Your Dataset

Open up slr05.xls and scan for extreme values in X or Y—outliers can cause massive loss spikes that break training. You can either remove these outliers or replace them with median values to stabilize training.

Quick Troubleshooting Order

Start with feature normalization → then lower learning rate/switch to Adam → if still broken, try gradient clipping or weight initialization tweaks. Most of the time, normalization alone fixes the NaN issue!

内容的提问来源于stack exchange,提问作者Jespar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:49:59