LSTM模型训练损失恒为零、精度异常问题排查求助
Hey there! Let's dig into why your LSTM is showing zero loss and poor accuracy—there are a few critical issues in your code that are causing this behavior:
1. Training/Test Set Division is Completely Broken
You're using int(0.3)*len(features_normalized) to split your data. Since int(0.3) evaluates to 0, your train_x and train_y are empty arrays! That means your model isn't actually training on any data at all—hence the zero loss every single epoch.
Fix:
Calculate 30% of the data correctly by moving the multiplication inside the int() call:
train_x = features_normalized[:int(0.3 * len(features_normalized))] train_y = labels[:int(0.3 * len(features_normalized))] test_x = features_normalized[-int(0.3 * len(features_normalized)):] test_y = labels[-int(0.3 * len(features_normalized)):]
2. Invalid Comparison Operators in Label Generation
You're using HTML escape characters > and < instead of Python's actual > and < operators. This will either throw a syntax error or fail to assign labels correctly—likely leaving most labels stuck as [0,1,0] since the equality check is the only one that works.
Fix:
Replace the escaped characters with proper Python operators:
for _ in range(len(df.index)): ret = np.zeros(3) if df['Percentage_Change'][_] > 0: ret[0] = 1 elif df['Percentage_Change'][_] == 0: ret[1] = 1 elif df['Percentage_Change'][_] < 0: ret[2] = 1 ret_ult.append(ret)
3. No Feature Normalization
Your features (like Adj. Close and Adj. Volume) have wildly different scales. LSTMs are extremely sensitive to unnormalized data, which can make training unstable or prevent the model from learning meaningful patterns entirely.
Fix:
Normalize your features to a standard range (e.g., 0-1) using scikit-learn's scaler:
from sklearn.preprocessing import MinMaxScaler scaler = MinMaxScaler(feature_range=(0,1)) features_normalized = scaler.fit_transform(np.array(df[:-30]))
4. Incorrect RNN Time Step Setup
You've set n_chunks=1 and chunk_size=6, which means each input sample is a single time step with 6 features. But LSTMs are built to process sequential data—you need to create sliding windows of past data to predict future labels.
Example Fix for Sliding Windows:
Create sequences of 30 past time steps to predict the next label:
def create_sequences(data, labels, seq_length): xs, ys = [], [] for i in range(len(data) - seq_length): xs.append(data[i:i+seq_length]) ys.append(labels[i+seq_length]) return np.array(xs), np.array(ys) # Use 70% for training, 30% for testing seq_length = 30 train_split = int(0.7 * len(features_normalized)) train_x, train_y = create_sequences(features_normalized[:train_split], labels[:train_split], seq_length) test_x, test_y = create_sequences(features_normalized[train_split:], labels[train_split:], seq_length) # Update placeholders and RNN parameters to match sequence length n_chunks = seq_length chunk_size = 6 x = tf.placeholder('float', [None, n_chunks, chunk_size])
5. Redundant Batch Loop Logic
Your training loop has a nested while loop inside a for loop, which will cause you to iterate over the training data multiple times per epoch. This is inefficient and can lead to unexpected training behavior.
Fix:
Simplify the loop to process batches correctly:
def train_neural_network(x): prediction = RNN_neural_network_model(x) cost = tf.reduce_mean(tf.nn.softmax_cross_entropy_with_logits(logits=prediction, labels=y)) optimizer = tf.train.AdamOptimizer(learning_rate=0.001).minimize(cost) with tf.Session() as sess: # Use the non-deprecated initializer sess.run(tf.global_variables_initializer()) for epoch in range(hm_epochs): epoch_loss = 0 i = 0 while i < len(train_x): start = i end = i + batch_size batch_x = train_x[start:end] batch_y = train_y[start:end] i += batch_size _, c = sess.run([optimizer, cost], feed_dict={x: batch_x, y: batch_y}) epoch_loss += c print('Epoch', epoch+1, 'completed out of', hm_epochs, 'loss', epoch_loss) correct = tf.equal(tf.argmax(prediction, 1), tf.argmax(y, 1)) accuracy = tf.reduce_mean(tf.cast(correct, 'float')) print('Accuracy:', accuracy.eval({x: test_x, y: test_y}))
Bonus: Deprecated TensorFlow Function
tf.initialize_all_variables() is outdated—switch to tf.global_variables_initializer() as shown in the fix above to avoid warnings.
After applying these fixes, your model should start training properly (you'll see loss decrease over epochs, and accuracy should improve as the model learns patterns in the sequential financial data).
内容的提问来源于stack exchange,提问作者Daniel Xin Li

