手动实现LSTM/RNN效果不佳 求问题排查与优化方案
Hey there! Let's break down your issues step by step—totally get how overwhelming manual LSTM implementations can be when you're new to deep learning. Let's tackle your core questions first, then dive into concrete fixes for your code.
Your model's poor performance is almost certainly due to critical implementation mistakes in the LSTM logic, data handling, and training pipeline. Here's the breakdown of key issues and fixes:
Critical Implementation Flaws
- Incorrect Weight Dimensions: Your gate weights (
w_igate,w_fgate, etc.) are defined with shape[sequence_length, hidden_layers_size], but this is wrong. For a univariate time series (each time step has 1 feature), the input-to-gate weight matrices should have shape[1, hidden_layers_size]—not the sequence length. The sequence length is the number of time steps you feed into the model, not the feature dimension. - Batch Processing Failure: Your code doesn't properly handle batch data. Variables like
ht_prevare initialized as(1, hidden_layers_size), but for a batch size greater than 1, this should be(batch_size, hidden_layers_size)to track state for each sample in the batch separately. Your list-based state variables (ft,it, etc.) also don't work with tensor operations for batches. - Missing Training Logic: There's no loss function definition or backpropagation code in your snippet. Without computing gradients and updating weights, your model isn't learning anything at all.
- Static Initial State: You're using fixed initial hidden/cell states (
ct_prev,ht_prev) that never reset or update between batches/epochs. This contaminates state across different batches and prevents proper learning. - Incomplete Data Preparation: Your
load_datafunction cuts off before definingdata_y(the target values for your many-to-one task). For many-to-one, each input sequence should map to the next value in the series (e.g., sequence[2,4,6]maps to8for your 2x sequence data).
Optimization Steps
- Fix Weight Shapes: Change all input-to-gate weights to
[1, hidden_layers_size](since your input is univariate). - Implement Batch-Aware State Handling: Use tensors instead of lists for LSTM states, and initialize hidden/cell states as
(batch_size, hidden_layers_size)zeros at the start of each batch. - Add Loss & Backpropagation: Use Mean Squared Error (MSE) as your loss function (ideal for regression tasks), and use TensorFlow's
GradientTapeto compute gradients and update weights with an optimizer like Adam. - Proper Data Preparation: Generate
data_xas(num_samples, sequence_length, 1)(adding the feature dimension) anddata_yas(num_samples, 1)(the target value for each sequence). - Tune Hyperparameters: Start with a small hidden layer size (e.g., 16 or 32), learning rate between
1e-3and1e-4, and enough epochs (50-200) to let the model converge. - Add Regularization: Optional but helpful—add L2 regularization to weights or dropout layers to prevent overfitting (though with your simple 2x sequence data, overfitting might not be the issue yet).
Your de_normalize function is conceptually correct, but you need to follow two key rules to get accurate original values:
- Use Training-Scaling Parameters: Always use the
data_maxanddata_mincomputed from your training data to de-normalize predictions—never use values from the test set. This ensures consistency between training and inference. - Verify the Formula: Your current function is correct if
m1 = data_maxandm2 = data_min:
If your model uses an activation function that outputs outsidedef de_normalize(self, value, data_max, data_min): return (value * (data_max - data_min)) + data_min[0,1](e.g., tanh), you'll need to first map the output back to[0,1]before applying this function. For your 2x sequence data, since you normalized to[0,1], using a sigmoid activation on the output layer is a good match.
dynamic_rnn Only Needs De-Normalization TensorFlow's dynamic_rnn (and the higher-level tf.keras.layers.LSTM) is a fully tested, optimized implementation that handles all the low-level details you're missing in your manual code:
- Correct weight dimensions and batch-aware state management
- Proper gradient computation and backpropagation through time (BPTT)
- Automatic handling of sequence lengths and tensor shapes
- Valid initialization of hidden/cell states
When you use dynamic_rnn correctly (with input shape (batch_size, sequence_length, input_dim)), the model's output directly corresponds to the normalized target values you trained on. De-normalizing these outputs gives you the original scale because the model learned to predict the normalized target accurately—something your manual implementation isn't doing right now.
Here's a revised version of key parts of your code to fix the issues above:
Fixed Initialization with Correct Weight Shapes
import tensorflow as tf from tensorflow.initializers import GlorotUniform # Replaces deprecated tf.contrib.xavier_initializer class lstm_network(): name = "lstm_" def __init__(self, config_params): self.sequence_length = config_params["sequence_length"] self.batch_size = config_params["batch_size"] self.hidden_layers_size = config_params["hidden_layers_size"] self.data_path = config_params["data_path"] self.n_epochs = config_params["no_of_epochs"] self.learning_rate = config_params["learning_rate"] # Correct weight shapes: input_dim=1 (univariate data) init = GlorotUniform() self.w_igate = tf.Variable(init(shape=[1, self.hidden_layers_size]), name='w_igate') self.w_fgate = tf.Variable(init(shape=[1, self.hidden_layers_size]), name='w_fgate') self.w_ogate = tf.Variable(init(shape=[1, self.hidden_layers_size]), name='w_ogate') self.w_cgate = tf.Variable(init(shape=[1, self.hidden_layers_size]), name='w_cgate') self.u_igate = tf.Variable(init(shape=[self.hidden_layers_size, self.hidden_layers_size]), name='u_igate') self.u_fgate = tf.Variable(init(shape=[self.hidden_layers_size, self.hidden_layers_size]), name='u_fgate') self.u_ogate = tf.Variable(init(shape=[self.hidden_layers_size, self.hidden_layers_size]), name='u_ogate') self.u_cgate = tf.Variable(init(shape=[self.hidden_layers_size, self.hidden_layers_size]), name='u_cgate') self.w_output_layer = tf.Variable(init(shape=[self.hidden_layers_size, 1]), name='w_output_layer') self.data_max = None self.data_min = None print("\n LSTM network initialized with correct configurations")
Complete Data Loading
def load_data(self): import csv import numpy as np with open(self.data_path, 'r') as data_file: data_reader = csv.reader(data_file, delimiter=',') self.data = [float(row[1]) for row in data_reader] self.data_max = float(max(self.data)) self.data_min = float(min(self.data)) # Normalize to [0,1] self.data = np.array([(x - self.data_min)/(self.data_max - self.data_min) for x in self.data]) # Generate many-to-one sequences self.data_x = [] self.data_y = [] for i in range(len(self.data) - self.sequence_length): self.data_x.append(self.data[i:i+self.sequence_length]) self.data_y.append(self.data[i+self.sequence_length]) # Reshape for TensorFlow: (num_samples, sequence_length, input_dim) self.data_x = np.array(self.data_x).reshape(-1, self.sequence_length, 1) self.data_y = np.array(self.data_y).reshape(-1, 1) # Split into train/test (e.g., 80/20 split) split_idx = int(0.8 * len(self.data_x)) self.train_x, self.test_x = self.data_x[:split_idx], self.data_x[split_idx:] self.train_y, self.test_y = self.data_y[:split_idx], self.data_y[split_idx:]
Training Loop with Backpropagation
def train(self): optimizer = tf.keras.optimizers.Adam(learning_rate=self.learning_rate) self.training_loss = [] for epoch in range(self.n_epochs): epoch_loss = 0.0 # Shuffle training data each epoch indices = np.random.permutation(len(self.train_x)) train_x_shuffled = self.train_x[indices] train_y_shuffled = self.train_y[indices] # Iterate over batches for i in range(0, len(train_x_shuffled), self.batch_size): x_batch = train_x_shuffled[i:i+self.batch_size] y_batch = train_y_shuffled[i:i+self.batch_size] # Handle last batch if it's smaller than batch_size if len(x_batch) < self.batch_size: continue with tf.GradientTape() as tape: # Initialize hidden/cell states for the batch h_prev = tf.zeros((self.batch_size, self.hidden_layers_size)) c_prev = tf.zeros((self.batch_size, self.hidden_layers_size)) # Forward pass through each time step for t in range(self.sequence_length): x_t = x_batch[:, t, :] # Shape: (batch_size, 1) # Compute gate values it = tf.sigmoid(tf.matmul(x_t, self.w_igate) + tf.matmul(h_prev, self.u_igate)) ft = tf.sigmoid(tf.matmul(x_t, self.w_fgate) + tf.matmul(h_prev, self.u_fgate)) ot = tf.sigmoid(tf.matmul(x_t, self.w_ogate) + tf.matmul(h_prev, self.u_ogate)) ct_tilde = tf.tanh(tf.matmul(x_t, self.w_cgate) + tf.matmul(h_prev, self.u_cgate)) # Update cell and hidden states c_prev = ft * c_prev + it * ct_tilde h_prev = ot * tf.tanh(c_prev) # Predict output from final hidden state pred = tf.matmul(h_prev, self.w_output_layer) # Compute MSE loss loss = tf.reduce_mean(tf.square(pred - y_batch)) # Compute gradients and update weights grads = tape.gradient(loss, [self.w_igate, self.w_fgate, self.w_ogate, self.w_cgate, self.u_igate, self.u_fgate, self.u_ogate, self.u_cgate, self.w_output_layer]) optimizer.apply_gradients(zip(grads, [self.w_igate, self.w_fgate, self.w_ogate, self.w_cgate, self.u_igate, self.u_fgate, self.u_ogate, self.u_cgate, self.w_output_layer])) epoch_loss += loss.numpy() avg_epoch_loss = epoch_loss / (len(train_x_shuffled) // self.batch_size) self.training_loss.append(avg_epoch_loss) print(f"Epoch {epoch+1}/{self.n_epochs}, Average Loss: {avg_epoch_loss:.6f}")
Testing & De-Normalization
def de_normalize(self, value): return (value * (self.data_max - self.data_min)) + self.data_min def test(self): test_loss = 0.0 predictions = [] true_values = [] for i in range(0, len(self.test_x), self.batch_size): x_batch = self.test_x[i:i+self.batch_size] y_batch = self.test_y[i:i+self.batch_size] if len(x_batch) < self.batch_size: continue h_prev = tf.zeros((self.batch_size, self.hidden_layers_size)) c_prev = tf.zeros((self.batch_size, self.hidden_layers_size)) for t in range(self.sequence_length): x_t = x_batch[:, t, :] it = tf.sigmoid(tf.matmul(x_t, self.w_igate) + tf.matmul(h_prev, self.u_igate)) ft = tf.sigmoid(tf.matmul(x_t, self.w_fgate) + tf.matmul(h_prev, self.u_fgate)) ot = tf.sigmoid(tf.matmul(x_t, self.w_ogate) + tf.matmul(h_prev, self.u_ogate)) ct_tilde = tf.tanh(tf.matmul(x_t, self.w_cgate) + tf.matmul(h_prev, self.u_cgate)) c_prev = ft * c_prev + it * ct_tilde h_prev = ot * tf.tanh(c_prev) pred = tf.matmul(h_prev, self.w_output_layer) loss = tf.reduce_mean(tf.square(pred - y_batch)) test_loss += loss.numpy() # De-normalize predictions and true values predictions.extend(self.de_normalize(pred.numpy()).flatten()) true_values.extend(self.de_normalize(y_batch.numpy()).flatten()) self.testing_loss = test_loss / (len(self.test_x) // self.batch_size) print(f"\nTesting Loss: {self.testing_loss:.6f}") return predictions, true_values
内容的提问来源于stack exchange,提问作者Rohan Baisantry

