You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中DNNLinearCombinedRegressor回归模型损失过高求助

Troubleshooting High MSE in TensorFlow DNNLinearCombinedRegressor for Regression Task

That’s definitely odd—with your target labels ranging from 0-40, an MSE over 1000 means your model isn’t learning anything meaningful from the data. Let’s break down the most likely issues and fixes step by step:

1. Critical Data Mix-Up (The Most Obvious Red Flag)

Looking at your code, you’re training the model on test_data instead of a dedicated training dataset! This is a huge mistake—if you’re using held-out test data for training, the model can’t generalize, and it’ll either overfit immediately or fail to learn patterns entirely.

Fix:
Replace test_data in the model.train call with your actual training dataset:

def main():
    shutil.rmtree(modeldir, ignore_errors=True)  # Add ignore_errors to avoid missing directory errors
    model = build_estimator(modeldir)
    # Use train_data instead of test_data here!
    model.train(input_function = lambda: input_function(200, train_data))

2. Input Function Issues

Your input function is the backbone of data feeding—if it’s mishandling features or labels, the model will never converge. Check these points:

  • Feature-Label Separation: Ensure your input_function correctly splits features and the target label (returning (features, labels) instead of mixing them together).
  • Normalize Continuous Features: DNNs are extremely sensitive to unnormalized numeric data. If you have continuous features (e.g., age, income), scale them to a small range (0-1 with MinMaxScaler or mean=0/variance=1 with StandardScaler).
  • Batch Size & Shuffling: Add shuffling for training data and set a reasonable batch size (64-128 works for most cases) to stabilize training.

Example Corrected Input Function:

from sklearn.preprocessing import MinMaxScaler

def input_function(num_epochs, data_set, is_training=True):
    # Split features and target label
    features = data_set.drop("your_target_column", axis=1)
    labels = data_set["your_target_column"]
    
    # Normalize continuous features
    continuous_cols = ["col1", "col2", "col3"]  # Replace with your actual continuous columns
    scaler = MinMaxScaler()
    features[continuous_cols] = scaler.fit_transform(features[continuous_cols])
    
    # Build TensorFlow Dataset
    dataset = tf.data.Dataset.from_tensor_slices((dict(features), labels))
    if is_training:
        dataset = dataset.shuffle(buffer_size=len(data_set)).repeat(num_epochs).batch(64)
    else:
        dataset = dataset.batch(64)
    return dataset

3. Feature Column Misconfiguration

Incorrect feature column setup is another common culprit for poor regression performance:

  • Wide Columns: For categorical features, use categorical_column_with_vocabulary_list (if you know all possible values) or categorical_column_with_hash_bucket (for high-cardinality features). Avoid adding continuous features to wide columns unless you’re binning them.
  • Deep Columns: Wrap continuous features in numeric_column, and convert categorical features to embedding_column (use a reasonable embedding dimension—min(600, round(category_cardinality**0.25)) is a safe rule of thumb).
  • No Label Leakage: Double-check that your target column isn’t accidentally included in wide_columns or deep_columns.

4. Model Architecture & Hyperparameters

Your current hidden units [1024, 512, 256] might be overkill for a target range of 0-40—too much model capacity can lead to overfitting or gradient instability. Try these adjustments:

  • Shrink DNN Size: Start with smaller hidden units like [256, 128, 64] and scale up only if needed.
  • Explicit Optimizers: Set custom optimizers with lower learning rates to stabilize training:
    return tf.estimator.DNNLinearCombinedRegressor(
        model_dir = model_dir,
        linear_feature_columns = wide_columns,
        dnn_feature_columns = deep_columns,
        dnn_hidden_units = [256, 128, 64],
        linear_optimizer=tf.keras.optimizers.Ftrl(learning_rate=0.01),
        dnn_optimizer=tf.keras.optimizers.Adam(learning_rate=0.001)
    )
    
  • Check Logs: Enable TensorFlow logging to track loss trends:
    tf.logging.set_verbosity(tf.logging.INFO)
    
    If loss stays flat from the start, the model isn’t learning—this usually points to input data or feature column issues.

5. Validate Target Data

Finally, confirm your target labels are clean:

  • No missing values or outliers (e.g., values outside 0-40)
  • Labels are numeric (not string-encoded)
  • No class imbalance (though less critical for regression)

内容的提问来源于stack exchange,提问作者guru

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:01:01