回归模型预测值全为负值求助:含18特征的NN网络结构说明
Hey there! Let’s work through why your neural network is only spitting out negative predictions for your regression task. I’ve gone through your model code and setup, and here are the most likely issues to fix:
1. Output Layer Activation (or Lack Thereof)
Your final layer is Dense(1) with no activation function, which defaults to a linear activation. That’s fine if your target values include both positive and negative numbers—but if your actual y_train/y_test have positive values, the model might be stuck in a local minimum where it only predicts negatives.
Fix:
- First, check your target distribution to confirm if positive values exist:
print(f"y_train min: {y_train.min()}, max: {y_train.max()}, mean: {y_train.mean()}") - If your targets are all positive, add a ReLU activation to the output layer to restrict predictions to non-negative values:
Corr = Dense(1, activation='relu')(X) - If targets include both positive and negative values, keep the linear activation but focus on other fixes below.
2. Missing Data Preprocessing
Unscaled features or labels can throw off your model’s training, especially with deep networks. If your features have wildly different scales, the model’s weights might update in a way that pushes predictions toward negative values.
Fix:
Normalize/standardize your features (and optionally your labels) using scikit-learn:
from sklearn.preprocessing import StandardScaler # Scale features scaler_X = StandardScaler() X_train_scaled = scaler_X.fit_transform(X_train) X_test_scaled = scaler_X.transform(X_test) # If your target values have a large range, scale them too (reverse after prediction) scaler_y = StandardScaler() y_train_scaled = scaler_y.fit_transform(y_train.reshape(-1, 1)) y_test_scaled = scaler_y.transform(y_test.reshape(-1, 1)) # Train on scaled data model.fit(X_train_scaled, y_train_scaled, validation_data=(X_test_scaled, y_test_scaled), epochs=50, batch_size=512, verbose=1) # Reverse scaling for predictions y_pred = scaler_y.inverse_transform(model.predict(X_test_scaled))
3. Suboptimal Training Setup
Your current training parameters might be preventing the model from converging properly:
- SGD Optimizer: SGD with the default learning rate (0.01) can be too aggressive or too slow for deep networks. Adam’s adaptive learning rate often works better for these cases.
- Too Few Epochs: 20 epochs might not be enough for your 6-layer network to learn the patterns in your data.
- Weight Initialization: LeakyReLU works best with He-normal initialization, which helps prevent vanishing gradients.
Fixes:
- Switch to Adam optimizer or adjust SGD’s learning rate:
from tensorflow.keras.optimizers import Adam, SGD # Option 1: Use Adam (recommended for deep networks) model.compile(optimizer=Adam(learning_rate=0.001), loss=tf.keras.losses.Huber(), metrics=['mse', 'mae']) # Option 2: Tune SGD with momentum model.compile(optimizer=SGD(learning_rate=0.001, momentum=0.9), loss=tf.keras.losses.Huber(), metrics=['mse', 'mae']) - Increase epochs to 50 or 100, and monitor training/validation loss to see if predictions start to include positives as training progresses.
- Add He-normal initialization to your Dense layers:
X = Dense(1024, kernel_initializer='he_normal')(features) # Repeat this for all subsequent Dense layers
4. Loss Function Check
Your custom Huber loss wrapper is okay, but in newer TensorFlow versions, it’s better to use the built-in loss class directly to avoid any signature issues:
model.compile(optimizer='Adam', loss=tf.keras.losses.Huber(), metrics=['mse', 'mae'])
Next Steps to Debug
- Start by checking your target variable’s distribution—this will immediately tell you if an output activation is needed.
- Apply feature scaling first, as this is the most common fix for unexpected prediction biases.
- Adjust your optimizer and training epochs if scaling alone doesn’t help.
内容的提问来源于stack exchange,提问作者Pocholo Mendiola

