基于TensorFlow的无测试集时间序列未来预测技术咨询
Absolutely! TensorFlow (especially with its Keras API and dedicated time series tools) is totally up to this task—you don’t need a separate test set if your sole goal is predicting the unseen 2019 and 2020 values. Let’s break down how to do this, with a hands-on example and key resources to deepen your learning.
This is a classic multi-step time series forecasting problem. Since you don’t need to evaluate model performance on a held-out test set (though I’d still recommend doing a quick validation check on 2018 data if possible), you can use your entire 2010-2018 dataset to train a model, then iteratively generate predictions for future time steps.
LSTMs are a go-to choice for time series because they capture long-term temporal patterns. Here’s a complete, runnable example:
Step 1: Import Dependencies
import tensorflow as tf from tensorflow.keras.models import Sequential from tensorflow.keras.layers import LSTM, Dense import numpy as np import pandas as pd from sklearn.preprocessing import MinMaxScaler
Step 2: Load and Preprocess Data
Replace the simulated data below with your actual 2010-2018 dataset:
# Simulate your dataset (replace this with pd.read_csv("your_data.csv")) dates = pd.date_range(start="2010-01-01", end="2018-12-31", freq="M") values = np.random.randn(len(dates)).cumsum() + 100 # Simulated trending data df = pd.DataFrame({"date": dates, "value": values}) # Convert time series to supervised learning format def create_sequences(data, look_back=12): """Use past `look_back` time steps to predict the next step""" X, y = [], [] for i in range(len(data) - look_back): X.append(data[i : i + look_back, 0]) y.append(data[i + look_back, 0]) return np.array(X), np.array(y) # Normalize data (critical for LSTM performance) scaler = MinMaxScaler(feature_range=(0, 1)) scaled_values = scaler.fit_transform(df["value"].values.reshape(-1, 1)) # Create training data (use all 2010-2018 data) look_back = 12 # Use past 12 months to predict the next month X_train, y_train = create_sequences(scaled_values, look_back) # Reshape input for LSTM: [samples, time steps, features] X_train = np.reshape(X_train, (X_train.shape[0], X_train.shape[1], 1))
Step 3: Build and Train the LSTM Model
# Define the LSTM model model = Sequential([ LSTM(50, return_sequences=True, input_shape=(look_back, 1)), LSTM(50), Dense(1) # Output layer for single value prediction ]) model.compile(loss="mean_squared_error", optimizer="adam") # Train the model model.fit(X_train, y_train, epochs=100, batch_size=32, verbose=1)
Step 4: Predict 2019 and 2020 Values
We’ll iteratively predict each future month, using the previous predictions as input for the next step:
# Predict 24 months (2019 + 2020) num_future_steps = 24 predictions = [] # Start with the last 12 months of training data current_sequence = scaled_values[-look_back:] for _ in range(num_future_steps): # Reshape for model input seq_reshaped = np.reshape(current_sequence, (1, look_back, 1)) # Predict next value next_pred = model.predict(seq_reshaped, verbose=0) predictions.append(next_pred[0][0]) # Update sequence: drop oldest value, add new prediction current_sequence = np.append(current_sequence[1:], next_pred, axis=0) # Convert predictions back to original scale predictions = scaler.inverse_transform(np.array(predictions).reshape(-1, 1)) # Create a dataframe for results future_dates = pd.date_range(start="2019-01-01", end="2020-12-31", freq="M") future_predictions = pd.DataFrame({ "date": future_dates, "predicted_value": predictions.flatten() }) print(future_predictions)
Even without external links, TensorFlow’s official docs have everything you need to dive deeper:
- Keras Time Series Guide: Explains LSTMs, TCNs (Temporal Convolutional Networks), and best practices for data preprocessing and forecasting.
- TensorFlow Time Series (TFTS): A dedicated module for time series tasks that handles auto-feature engineering and rolling predictions out of the box.
- Keras Example Library: Contains pre-built notebooks for multi-step forecasting, seasonal prediction, and even transformer-based time series models.
- If your data has strong seasonal patterns, add seasonal features (like month-of-year encoding) to improve predictions.
- Even without a formal test set, do a quick validation: train on 2010-2017 data, predict 2018, and compare to actual 2018 values to gauge model accuracy.
- For better performance with long sequences, consider using Transformer models (e.g.,
TimeSeriesTransformerin Keras).
内容的提问来源于stack exchange,提问作者Clinton Prakash

