TensorFlow中玩具级LSTM模型训练失败技术求助
Let's break down the key issues in your setup and walk through the fixes to get your toy model training successfully.
1. Critical Data Issues
a. Unshuffled Data
Your current data is ordered: all positive samples first, then all negative samples. This makes the model learn to predict based on the sample's position in the dataset, not the sequence content itself. You must shuffle the data and labels together before training.
Fix code:
# Shuffle data and labels (use sklearn or manual shuffling) from sklearn.utils import shuffle data, labels = shuffle(data, labels) # Or manual shuffling if you don't want to use sklearn: indices = np.random.permutation(len(data)) data = data[indices] labels = labels[indices]
b. Incorrect Input Shape for LSTM
LSTMs in TensorFlow expect 3D input: (batch_size, time_steps, features). Your current data is 2D ((1000, 5)), where each sequence has 5 time steps, each with 1 feature. Reshape it to 3D:
data = data.reshape((data.shape[0], data.shape[1], 1))
Update your placeholder to match this shape:
_, x_width, num_features = data.shape X = tf.placeholder(tf.float32, [None, x_width, num_features])
2. Model Architecture Fixes
Your model snippet was cut off, so here's a complete, minimal LSTM setup tailored to your problem:
import tensorflow as tf # Hyperparameters (adjust as needed) num_hidden = 16 # Small enough for a toy problem num_classes = 2 learning_rate = 0.001 epochs = 20 batch_size = 32 # Placeholders Y = tf.placeholder(tf.float32, [None, num_classes]) # LSTM Layer lstm_cell = tf.nn.rnn_cell.LSTMCell(num_hidden, forget_bias=1.0) outputs, _ = tf.nn.dynamic_rnn(lstm_cell, X, dtype=tf.float32) # Use the last output of the LSTM for sequence classification last_sequence_output = outputs[:, -1, :] # Output layer with softmax for one-hot labels logits = tf.layers.dense(last_sequence_output, num_classes) prediction = tf.nn.softmax(logits) # Loss and optimizer (match your one-hot labels) loss = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits_v2(logits=logits, labels=Y)) optimizer = tf.train.AdamOptimizer(learning_rate).minimize(loss) # Accuracy metric correct_pred = tf.equal(tf.argmax(prediction, 1), tf.argmax(Y, 1)) accuracy = tf.reduce_mean(tf.cast(correct_pred, tf.float32))
Key points here:
- We take the last output of the LSTM sequence—this captures the final state needed to classify the sequence.
- Using
softmax_cross_entropy_with_logits_v2matches your one-hot label format. - The Adam optimizer is a safe default for most cases.
3. Training Loop Best Practices
Make sure your training loop handles batching and runs the initializer:
# Initialize variables init = tf.global_variables_initializer() with tf.Session() as sess: sess.run(init) for epoch in range(epochs): # Shuffle data each epoch for better training data_shuffled, labels_shuffled = shuffle(data, labels) # Iterate over batches for i in range(0, len(data), batch_size): batch_x = data_shuffled[i:i+batch_size] batch_y = labels_shuffled[i:i+batch_size] # Run training step _, batch_loss, batch_acc = sess.run([optimizer, loss, accuracy], feed_dict={X: batch_x, Y: batch_y}) # Print progress each epoch print(f"Epoch {epoch+1}/{epochs}, Loss: {batch_loss:.4f}, Accuracy: {batch_acc:.4f}")
Why This Works
Your sequence difference is tiny (only the last element changes: 5 vs 6), but the LSTM will learn to focus on that final element once the data is shuffled and the input shape is correct. With this setup, you should see accuracy reach 100% within a few epochs.
内容的提问来源于stack exchange,提问作者Gerry

