You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Keras批量训练在线预测无学习效果问题求助

诊断你的Keras共享单车预测模型问题

Hey there, I’ve worked on several bike-sharing availability forecasting projects before, so let’s dig into why your model might be failing to learn. Most time-series model issues boil down to data prep, input structure, feature choices, or model setup—let’s break this down step by step.

1. First Check: Data Preprocessing & Scaling

Mixed feature scales are a common killer for neural networks. Your features have wildly different ranges:

  • Bike count might be 0–50 (or higher)
  • Day of year is 0–365
  • Day of week is 0–6
  • Hour of day is 0–23

If you’re feeding raw values into Keras without normalization, the model will prioritize features with larger scales (like day of year) and ignore smaller ones. Try these fixes:

  • Use StandardScaler or MinMaxScaler to normalize numerical features to a 0–1 or -1–1 range.
  • Critical: Fit the scaler only on your training data, then apply it to validation/test/online prediction data—never refit on new data, this will leak information.
  • Clean up outliers: If a station suddenly shows 100 bikes (way above its capacity), clip those values or mark them as missing and impute with median/mean.

2. Fix Your Time-Series Input Shape

Keras recurrent layers (LSTM/GRU) expect inputs in the shape (number_of_samples, time_steps, number_of_features). If you’re feeding single timesteps or flattening sequences, the model can’t learn temporal patterns. For your setup:

  • Each training sample should be a window of N past time steps, including all your features (bike count at each step, day of year, day of week, hour)
  • The target is the bike count at the next time step (or whatever prediction horizon you’re targeting)
  • Example: If N=24 (24 hours of history), each sample is (24, 4) (24 steps, 4 features)
  • For online prediction, maintain a sliding window of the last N time steps—make sure this window uses the exact same preprocessing as training data.

3. Improve Feature Engineering for Cyclical Variables

Raw integer values for cyclical features (hour, day of week, day of year) confuse the model. For example, hour 23 and hour 0 are adjacent in time, but as integers, their distance is 23. Fix this by encoding them as sine/cosine pairs:

import numpy as np

def encode_cyclical_feature(value, max_value):
    sin_val = np.sin(2 * np.pi * value / max_value)
    cos_val = np.cos(2 * np.pi * value / max_value)
    return sin_val, cos_val

# Example for hour (0-23)
hour_sin, hour_cos = encode_cyclical_feature(hour_data, 23)
# Repeat for day of week (max=6) and day of year (max=365)

This converts each cyclical feature into two continuous values that capture the circular nature of time—your model will pick up on daily/weekly/yearly patterns much easier.

4. Choose the Right Model Architecture

A simple Dense-only model can’t capture temporal dependencies in time-series data. Switch to a recurrent or temporal-focused model:

  • LSTM/GRU: Standard for time-series forecasting. Start with a small stack: LSTM(64, return_sequences=False, input_shape=(N, num_features)) followed by a Dense output layer (since this is a regression task).
  • TCN (Temporal Convolutional Network): Great for capturing long-range dependencies without vanishing gradient issues—Keras has easy-to-use TCN implementations.
  • Avoid overcomplicating: Start with a simple model first, then add dropout or additional layers if you see signs of overfitting.

5. Tune Training Hyperparameters

Even with good data and architecture, poor training settings can stop the model from learning:

  • Loss function: Use MeanSquaredError() or MeanAbsoluteError() for regression (bike count prediction is a regression task, not classification).
  • Optimizer: Start with Adam with a moderate learning rate (0.001 is a safe default). If the loss isn’t decreasing, try lowering it to 0.0001.
  • Batch size: For time-series, smaller batches (32–64) often work better than large ones, as they help the model learn temporal patterns faster.
  • Epochs: Use early stopping with a validation set—stop training if validation loss stops decreasing for 5–10 epochs to avoid overfitting.

6. Validate with Interpretable Metrics

Make sure you’re using metrics that matter for your use case:

  • MAE (Mean Absolute Error): Tells you the average number of bikes your prediction is off by—super interpretable for this task.
  • RMSE (Root Mean Squared Error): Penalizes larger errors more, which is useful if big misses are costly for your application.

Try these steps one by one—start with data scaling and input shape, since those are the most common culprits. Once you get the model learning on a single station, you can expand to batch training across multiple stations (maybe using a shared base model with station-specific heads, but that’s a next step).

内容的提问来源于stack exchange,提问作者user4296481

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:51:28