You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用TensorFlow Dataset API训练DNNRegressor时损失未稳定下降问题咨询

Hey there, let's break down your problem step by step—training loss not stabilizing with tf.estimator.DNNRegressor and the Dataset API is a common pain point, but we can work through it with targeted fixes and best practices.

1. Why Your Training Loss Might Be Misbehaving

Let’s start with the most likely culprits for unstable loss:

  • Feature Scaling is Non-Negotiable for DNNs
    DNNs are terrible at handling features with wildly different scales. Your continuous features (precipitation, temperature) might range from, say, 0-100 and -10-30, while categorical integers (weekday 0-6, season 0-n) are on a tiny scale. If you feed raw values, the model will waste cycles adjusting weights to compensate for scale differences instead of learning patterns. Normalize continuous features to mean 0 / standard deviation 1, and one-hot encode categorical features so each category gets its own binary input.

  • No Train/Eval Split
    If you’re training on all 1200 rows without holding out a validation set, you can’t tell if the loss instability is due to overfitting or poor learning. Split your data into 80% training and 20% validation—this lets you compare training loss vs validation loss to diagnose issues (e.g., if training loss drops but validation loss rises, you’re overfitting).

  • Dataset Shuffling & Batching Missteps

    • Shuffling buffer size: If you use a buffer smaller than your dataset (e.g., 100 for 1200 rows), the shuffle won’t be thorough, leading to biased gradient updates. Use buffer_size=len(data_df) to shuffle all rows each epoch.
    • Batch size: Too small (like 8) causes noisy gradients; too large (like 512) makes updates slow and can trap the model in local minima. For 1200 rows, try 32, 64, or 128.
    • Missing repeat(): If you don’t call repeat(num_epochs), the dataset will only iterate once, so the model stops training even if you set a high max_steps.
  • Misconfigured Model Hyperparameters

    • Learning rate: If loss bounces up and down a lot, your learning rate is too high (try dropping from 0.01 to 0.001). If loss decreases at a snail’s pace, nudge it up. Adam optimizer is usually safer than SGD here.
    • Hidden layers: A model that’s too shallow (e.g., 1 layer with 16 units) might underfit; too deep (3+ layers with 256 units) might overfit. Start with 2 layers of 64 and 32 units, then adjust.
2. Dataset API Best Practices for Your Pandas DataFrame

Since you’re working with a pandas DataFrame, here’s how to properly convert it to a Dataset for tf.estimator:

First, split your data and separate features/labels:

import tensorflow as tf
import pandas as pd
from sklearn.model_selection import train_test_split

# Load your data
df = pd.read_csv("your_data.csv")

# Split into train/validation sets
train_df, eval_df = train_test_split(df, test_size=0.2, random_state=42)

# Separate features and target label
train_features = train_df.drop("visitors", axis=1)
train_labels = train_df["visitors"]
eval_features = eval_df.drop("visitors", axis=1)
eval_labels = eval_df["visitors"]

Next, create a reusable input function—this avoids repetition and ensures consistency between train/eval:

def make_input_fn(data_df, label_df, num_epochs=10, shuffle=True, batch_size=64):
    def input_fn():
        # Convert pandas data to Dataset: features as dict, labels as tensor
        ds = tf.data.Dataset.from_tensor_slices((dict(data_df), label_df))
        if shuffle:
            ds = ds.shuffle(buffer_size=len(data_df))  # Shuffle all data each epoch
        # Repeat for specified epochs and batch
        ds = ds.repeat(num_epochs).batch(batch_size)
        return ds
    return input_fn

# Create train/eval input functions
train_input_fn = make_input_fn(train_features, train_labels, num_epochs=50, shuffle=True)
eval_input_fn = make_input_fn(eval_features, eval_labels, num_epochs=1, shuffle=False)
3. Correct Feature Column Setup for tf.estimator

tf.estimator relies on feature columns to understand your data—this is a common point of confusion. Here’s how to define them for your features:

# Normalize continuous features
def normalize_numeric_column(col_name):
    mean = train_features[col_name].mean()
    std = train_features[col_name].std()
    return tf.feature_column.numeric_column(
        col_name,
        normalizer_fn=lambda x: (x - mean) / std
    )

precipitation = normalize_numeric_column("precipitation")
temperature = normalize_numeric_column("temperature")

# One-hot encode categorical features
weekday = tf.feature_column.categorical_column_with_vocabulary_list(
    "weekday", vocabulary_list=[0, 1, 2, 3, 4, 5, 6]
)
weekday_one_hot = tf.feature_column.indicator_column(weekday)

season = tf.feature_column.categorical_column_with_vocabulary_list(
    "season", vocabulary_list=sorted(train_features["season"].unique())
)
season_one_hot = tf.feature_column.indicator_column(season)

week_number = tf.feature_column.categorical_column_with_vocabulary_list(
    "week_number", vocabulary_list=sorted(train_features["week_number"].unique())
)
week_number_one_hot = tf.feature_column.indicator_column(week_number)

# Combine all feature columns
feature_columns = [precipitation, temperature, weekday_one_hot, season_one_hot, week_number_one_hot]
4. Putting It All Together & Troubleshooting

Define and train your model:

model = tf.estimator.DNNRegressor(
    feature_columns=feature_columns,
    hidden_units=[64, 32],  # Start with this, adjust later
    optimizer=tf.optimizers.Adam(learning_rate=0.001),
    model_dir="./visitor_prediction_model",  # Saves checkpoints for later use
    dropout=0.2  # Optional: adds dropout to prevent overfitting
)

# Train the model
model.train(input_fn=train_input_fn)

# Evaluate performance
eval_results = model.evaluate(input_fn=eval_input_fn)
print(f"Validation Loss: {eval_results['average_loss']:.2f}")

If loss is still unstable after this:

  • Increase the number of epochs (try 50-100 for your 1200-row dataset)
  • Tune hidden units (e.g., try [32, 32] or [128, 64])
  • Add L2 regularization to the optimizer: optimizer=tf.optimizers.Adam(learning_rate=0.001, weight_decay=0.001)

内容的提问来源于stack exchange,提问作者tensormit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:42:01