You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

TensorFlow中玩具级LSTM模型训练失败技术求助

Fixing Your LSTM Training Failure for Sequence Classification

Let's break down the key issues in your setup and walk through the fixes to get your toy model training successfully.

1. Critical Data Issues

a. Unshuffled Data

Your current data is ordered: all positive samples first, then all negative samples. This makes the model learn to predict based on the sample's position in the dataset, not the sequence content itself. You must shuffle the data and labels together before training.

Fix code:

# Shuffle data and labels (use sklearn or manual shuffling)
from sklearn.utils import shuffle
data, labels = shuffle(data, labels)

# Or manual shuffling if you don't want to use sklearn:
indices = np.random.permutation(len(data))
data = data[indices]
labels = labels[indices]

b. Incorrect Input Shape for LSTM

LSTMs in TensorFlow expect 3D input: (batch_size, time_steps, features). Your current data is 2D ((1000, 5)), where each sequence has 5 time steps, each with 1 feature. Reshape it to 3D:

data = data.reshape((data.shape[0], data.shape[1], 1))

Update your placeholder to match this shape:

_, x_width, num_features = data.shape
X = tf.placeholder(tf.float32, [None, x_width, num_features])

2. Model Architecture Fixes

Your model snippet was cut off, so here's a complete, minimal LSTM setup tailored to your problem:

import tensorflow as tf

# Hyperparameters (adjust as needed)
num_hidden = 16  # Small enough for a toy problem
num_classes = 2
learning_rate = 0.001
epochs = 20
batch_size = 32

# Placeholders
Y = tf.placeholder(tf.float32, [None, num_classes])

# LSTM Layer
lstm_cell = tf.nn.rnn_cell.LSTMCell(num_hidden, forget_bias=1.0)
outputs, _ = tf.nn.dynamic_rnn(lstm_cell, X, dtype=tf.float32)

# Use the last output of the LSTM for sequence classification
last_sequence_output = outputs[:, -1, :]

# Output layer with softmax for one-hot labels
logits = tf.layers.dense(last_sequence_output, num_classes)
prediction = tf.nn.softmax(logits)

# Loss and optimizer (match your one-hot labels)
loss = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits_v2(logits=logits, labels=Y))
optimizer = tf.train.AdamOptimizer(learning_rate).minimize(loss)

# Accuracy metric
correct_pred = tf.equal(tf.argmax(prediction, 1), tf.argmax(Y, 1))
accuracy = tf.reduce_mean(tf.cast(correct_pred, tf.float32))

Key points here:

  • We take the last output of the LSTM sequence—this captures the final state needed to classify the sequence.
  • Using softmax_cross_entropy_with_logits_v2 matches your one-hot label format.
  • The Adam optimizer is a safe default for most cases.

3. Training Loop Best Practices

Make sure your training loop handles batching and runs the initializer:

# Initialize variables
init = tf.global_variables_initializer()

with tf.Session() as sess:
    sess.run(init)
    
    for epoch in range(epochs):
        # Shuffle data each epoch for better training
        data_shuffled, labels_shuffled = shuffle(data, labels)
        
        # Iterate over batches
        for i in range(0, len(data), batch_size):
            batch_x = data_shuffled[i:i+batch_size]
            batch_y = labels_shuffled[i:i+batch_size]
            
            # Run training step
            _, batch_loss, batch_acc = sess.run([optimizer, loss, accuracy], feed_dict={X: batch_x, Y: batch_y})
        
        # Print progress each epoch
        print(f"Epoch {epoch+1}/{epochs}, Loss: {batch_loss:.4f}, Accuracy: {batch_acc:.4f}")

Why This Works

Your sequence difference is tiny (only the last element changes: 5 vs 6), but the LSTM will learn to focus on that final element once the data is shuffled and the input shape is correct. With this setup, you should see accuracy reach 100% within a few epochs.

内容的提问来源于stack exchange,提问作者Gerry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:50:32