基于10年时序数据的股价预测模型选型及LSTM实现资源咨询
Hey there! Let's break down the best models for your stock price forecasting task (using 10 years of historical data to predict 1 month or 1 year ahead), plus a full LSTM implementation that digs into the core logic instead of just calling a library.
Here are the top options tailored to your use case:
Long Short-Term Memory (LSTM)
This is a go-to for time-series forecasting, especially with long historical data like your 10-year dataset. LSTMs solve the gradient vanishing problem of vanilla RNNs, so they can capture long-term dependencies in stock prices (like how events from 6 months ago might influence current trends). They’re flexible enough to handle both short-term (1 month) and longer-term (1 year) predictions if tuned properly.Improved Linear Regression (Time-Series Aware)
Plain vanilla linear regression won’t work well on its own—it doesn’t account for the sequential nature of stock data. But if you build a time-series linear model by adding relevant features like:- Sliding window metrics (past 30/60-day average price, volatility)
- Technical indicators (MACD, RSI, moving averages)
- Encoded time features (month, quarter, holiday flags)
It becomes a great baseline model. It’s fast, easy to interpret, and helps you measure how much value more complex models add.
Temporal Fusion Transformers (TFT)
If you have access to auxiliary data (like trading volume, macroeconomic indicators, news sentiment), TFTs are excellent. They’re designed for multi-variate time-series forecasting, handle long horizons (1 year) well, and provide built-in feature importance explanations—super useful for understanding what’s driving your predictions.Prophet
Facebook’s Prophet is perfect if you want a low-code, robust baseline. It excels at capturing trends and seasonality (like annual stock cycles) and handles outliers/holidays automatically. It’s great for quick experiments, even if you’re not a deep learning expert.Gated Recurrent Units (GRU)
A lighter alternative to LSTMs. GRUs combine the input and forget gates into a single update gate, making them faster to train with similar performance for many time-series tasks. If you’re working with limited computing resources, GRUs are a smart choice.
Below is a pure Python implementation of an LSTM, including all gate mechanisms, forward/backward propagation, and training logic—no shortcut library calls. This lets you see exactly how the model works under the hood:
import numpy as np from sklearn.preprocessing import MinMaxScaler # Data preprocessing: Normalization + sliding window creation def preprocess_data(data, seq_len, pred_len): # Normalize data to [0,1] for stable training scaler = MinMaxScaler(feature_range=(0, 1)) scaled_data = scaler.fit_transform(data.reshape(-1, 1)) X, y = [], [] # Create input sequences (past seq_len days) and target (future pred_len days) for i in range(seq_len, len(scaled_data) - pred_len + 1): X.append(scaled_data[i-seq_len:i, 0]) y.append(scaled_data[i:i+pred_len, 0]) return np.array(X), np.array(y), scaler # LSTM Unit: Implements full gate logic (input, forget, output, cell state) class LSTMUnit: def __init__(self, input_dim, hidden_dim): self.input_dim = input_dim self.hidden_dim = hidden_dim # Initialize weights with small random values to avoid saturation self.W_i = np.random.randn(input_dim + hidden_dim, hidden_dim) * 0.01 self.b_i = np.zeros((1, hidden_dim)) self.W_f = np.random.randn(input_dim + hidden_dim, hidden_dim) * 0.01 self.b_f = np.zeros((1, hidden_dim)) self.W_c = np.random.randn(input_dim + hidden_dim, hidden_dim) * 0.01 self.b_c = np.zeros((1, hidden_dim)) self.W_o = np.random.randn(input_dim + hidden_dim, hidden_dim) * 0.01 self.b_o = np.zeros((1, hidden_dim)) # Cache for backpropagation self.cache = {} def sigmoid(self, x): return 1 / (1 + np.exp(-x)) def tanh(self, x): return np.tanh(x) def forward(self, x_t, h_prev, c_prev): # Concatenate current input and previous hidden state concat = np.concatenate((x_t, h_prev), axis=1) # Input gate: Decide which new info to store in cell state i_t = self.sigmoid(np.dot(concat, self.W_i) + self.b_i) # Forget gate: Decide which old info to discard from cell state f_t = self.sigmoid(np.dot(concat, self.W_f) + self.b_f) # Candidate cell state: Generate new potential info c_tilde = self.tanh(np.dot(concat, self.W_c) + self.b_c) # Update cell state c_t = f_t * c_prev + i_t * c_tilde # Output gate: Decide which info to output as hidden state o_t = self.sigmoid(np.dot(concat, self.W_o) + self.b_o) # Current hidden state h_t = o_t * self.tanh(c_t) # Save intermediate values for backprop self.cache = { 'x_t': x_t, 'h_prev': h_prev, 'c_prev': c_prev, 'concat': concat, 'i_t': i_t, 'f_t': f_t, 'c_tilde': c_tilde, 'c_t': c_t, 'o_t': o_t, 'h_t': h_t } return h_t, c_t def backward(self, dh_t, dc_t): # Retrieve cached values from forward pass x_t, h_prev, c_prev = self.cache['x_t'], self.cache['h_prev'], self.cache['c_prev'] concat, i_t, f_t, c_tilde, c_t, o_t = self.cache['concat'], self.cache['i_t'], self.cache['f_t'], self.cache['c_tilde'], self.cache['c_t'], self.cache['o_t'] # Calculate gradients for output gate do_t = dh_t * self.tanh(c_t) do_t *= o_t * (1 - o_t) # Sigmoid derivative dW_o = np.dot(concat.T, do_t) db_o = np.sum(do_t, axis=0, keepdims=True) # Calculate gradients for cell state dc_t = dc_t + (dh_t * o_t) * (1 - np.square(self.tanh(c_t))) # Tanh derivative dc_tilde = dc_t * i_t * (1 - np.square(c_tilde)) dW_c = np.dot(concat.T, dc_tilde) db_c = np.sum(dc_tilde, axis=0, keepdims=True) # Calculate gradients for forget gate df_t = dc_t * c_prev * f_t * (1 - f_t) dW_f = np.dot(concat.T, df_t) db_f = np.sum(df_t, axis=0, keepdims=True) # Calculate gradients for input gate di_t = dc_t * c_tilde * i_t * (1 - i_t) dW_i = np.dot(concat.T, di_t) db_i = np.sum(di_t, axis=0, keepdims=True) # Calculate gradients for input and previous hidden state dconcat = (np.dot(do_t, self.W_o.T) + np.dot(dc_tilde, self.W_c.T) + np.dot(df_t, self.W_f.T) + np.dot(di_t, self.W_i.T)) dx_t = dconcat[:, :self.input_dim] dh_prev = dconcat[:, self.input_dim:] dc_prev = dc_t * f_t # Compile all gradients grads = { 'W_i': dW_i, 'b_i': db_i, 'W_f': dW_f, 'b_f': db_f, 'W_c': dW_c, 'b_c': db_c, 'W_o': dW_o, 'b_o': db_o, 'dx_t': dx_t, 'dh_prev': dh_prev, 'dc_prev': dc_prev } return grads # Full LSTM Model (single layer, easily extendable to multi-layer) class LSTMModel: def __init__(self, input_dim, hidden_dim, output_dim): self.input_dim = input_dim self.hidden_dim = hidden_dim self.output_dim = output_dim # Initialize LSTM unit self.lstm_unit = LSTMUnit(input_dim, hidden_dim) # Output layer weights self.W_out = np.random.randn(hidden_dim, output_dim) * 0.01 self.b_out = np.zeros((1, output_dim)) # Cache for forward pass self.cache = {} def forward(self, X): # X shape: (batch_size, sequence_length, input_dim) batch_size, seq_len, _ = X.shape h_prev = np.zeros((batch_size, self.hidden_dim)) c_prev = np.zeros((batch_size, self.hidden_dim)) # Store hidden states for each time step h_states = [] for t in range(seq_len): x_t = X[:, t, :] h_t, c_t = self.lstm_unit.forward(x_t, h_prev, c_prev) h_states.append(h_t) h_prev, c_prev = h_t, c_t # Use final hidden state to predict future values h_final = h_states[-1] y_pred = np.dot(h_final, self.W_out) + self.b_out self.cache = { 'X': X, 'h_states': h_states, 'h_final': h_final, 'y_pred': y_pred } return y_pred def backward(self, y_true): y_pred = self.cache['y_pred'] h_final = self.cache['h_final'] h_states = self.cache['h_states'] X = self.cache['X'] batch_size, seq_len, _ = X.shape # Output layer gradients dy = y_pred - y_true dW_out = np.dot(h_final.T, dy) db_out = np.sum(dy, axis=0, keepdims=True) dh_final = np.dot(dy, self.W_out.T) # Backpropagate through LSTM time steps (reverse order) dh_t = dh_final dc_t = np.zeros((batch_size, self.hidden_dim)) total_grads = { 'W_i': 0, 'b_i': 0, 'W_f': 0, 'b_f': 0, 'W_c': 0, 'b_c': 0, 'W_o': 0, 'b_o': 0, 'W_out': dW_out, 'b_out': db_out } for t in reversed(range(seq_len)): x_t = X[:, t, :] h_prev = h_states[t-1] if t > 0 else np.zeros((batch_size, self.hidden_dim)) c_prev = np.zeros((batch_size, self.hidden_dim)) if t == 0 else h_states[t-1] # Refresh LSTM unit cache for current time step self.lstm_unit.forward(x_t, h_prev, c_prev) grads = self.lstm_unit.backward(dh_t, dc_t) # Accumulate gradients across time steps for key in ['W_i', 'b_i', 'W_f', 'b_f', 'W_c', 'b_c', 'W_o', 'b_o']: total_grads[key] += grads[key] dh_t = grads['dh_prev'] dc_t = grads['dc_prev'] return total_grads def update_weights(self, grads, learning_rate): # Update LSTM unit weights self.lstm_unit.W_i -= learning_rate * grads['W_i'] self.lstm_unit.b_i -= learning_rate * grads['b_i'] self.lstm_unit.W_f -= learning_rate * grads['W_f'] self.lstm_unit.b_f -= learning_rate * grads['b_f'] self.lstm_unit.W_c -= learning_rate * grads['W_c'] self.lstm_unit.b_c -= learning_rate * grads['b_c'] self.lstm_unit.W_o -= learning_rate * grads['W_o'] self.lstm_unit.b_o -= learning_rate * grads['b_o'] # Update output layer weights self.W_out -= learning_rate * grads['W_out'] self.b_out -= learning_rate * grads['b_out'] # Training loop def train(model, X_train, y_train, epochs, learning_rate): # Reshape input to (batch_size, seq_len, input_dim) X_train = X_train.reshape(X_train.shape[0], X_train.shape[1], 1) for epoch in range(epochs): y_pred = model.forward(X_train) # Mean Squared Error for regression task loss = np.mean(np.square(y_pred - y_train)) grads = model.backward(y_train) model.update_weights(grads, learning_rate) # Print progress every 100 epochs if epoch % 100 == 0: print(f"Epoch {epoch:4d} | Loss: {loss:.6f}") # Prediction function def predict(model, X_test, scaler): X_test = X_test.reshape(X_test.shape[0], X_test.shape[1], 1) y_pred_scaled = model.forward(X_test) # Inverse normalize to get actual price values y_pred = scaler.inverse_transform(y_pred_scaled) return y_pred # Example usage (replace with your real stock price data) if __name__ == "__main__": # Simulate 10 years of daily stock prices (≈2500 trading days) stock_prices = np.random.randn(2500) * 10 + 100 # Hyperparameters seq_len = 60 # Use past 60 days to predict pred_len = 30 # Predict next 30 days (1 month) hidden_dim = 64 # Size of LSTM hidden state epochs = 1000 # Training cycles learning_rate = 0.001 # Step size for weight updates # Preprocess data X, y, scaler = preprocess_data(stock_prices, seq_len, pred_len) # Split into train/test sets (80/20 split) split_idx = int(0.8 * len(X)) X_train, X_test = X[:split_idx], X[split_idx:] y_train, y_test = y[:split_idx], y[split_idx:] # Initialize model model = LSTMModel(input_dim=1, hidden_dim=hidden_dim, output_dim=pred_len) # Train the model train(model, X_train, y_train, epochs, learning_rate) # Make predictions y_pred = predict(model, X_test, scaler) print("\nSample Predictions (next 30 days):") print(y_pred[:1]) # Print first prediction sequence
Notes on the LSTM Implementation
This code includes every core component of an LSTM:
- All four gate mechanisms (input, forget, output, cell update)
- Full forward and backward propagation with gradient calculations
- Data normalization and sliding window preparation
- Training loop with weight updates
You can extend it to multi-layer LSTMs by adding more LSTMUnit instances in the LSTMModel class, or add dropout for regularization if you overfit.
- For short-term (1 month) predictions: Start with LSTM or GRU—they handle short-term dependencies well.
- For long-term (1 year) predictions: Try TFT or Prophet (if seasonality is strong).
- Always use an improved linear model as a baseline to validate that complex models are actually adding value.
内容的提问来源于stack exchange,提问作者valer

