基于LSTM的RNN需3D输入?Keras搭建遇维度不匹配错误求助
Hey there, let's break down why you're getting this error and fix it properly.
First, the root cause: Keras LSTM layers expect a 3-dimensional input tensor with shape (number_of_samples, timesteps, number_of_features), but your train_spec array is currently 2D ((1415684, 1)). It's missing the critical timesteps dimension that LSTMs need to process sequential data.
Looking at your dataset, each patient has multiple time-stamped spectrum readings—this is exactly the kind of sequential data LSTMs excel at! Let's cover two possible scenarios based on your actual use case:
Scenario 1: Use each patient's sequential spectrum data as a sample (Recommended)
This is the right approach for LSTMs, since you'll leverage the temporal patterns in each patient's data over time. Here's how to adjust your code:
- Group data by patient ID to create sequence samples
- Standardize sequence lengths (since patients may have different numbers of readings)
- Reshape data to match LSTM's 3D input requirement
import numpy as np import pandas as pd from keras.models import Sequential from keras.layers import Dense, Dropout, Activation, LSTM from keras.preprocessing.sequence import pad_sequences # Load data with pandas for easier grouping train_df = pd.read_csv("TrainDatasetFinal.txt", header=None, names=["patient_id", "time", "x", "y", "z", "amplitude", "spectrum", "label"]) test_df = pd.read_csv("testDatasetFinal.txt", header=None, names=["patient_id", "time", "x", "y", "z", "amplitude", "spectrum", "label"]) # Function to convert raw data into sequence samples def prepare_sequences(df): sequences = [] labels = [] # Group data by each patient for patient_id, group in df.groupby("patient_id"): # Extract spectrum values as a sequence (1 feature per timestep) spectrum_seq = group["spectrum"].values.reshape(-1, 1) sequences.append(spectrum_seq) # Each patient has a single label (0/1), grab the first one labels.append(group["label"].iloc[0]) # Pad/truncate sequences to a uniform length max_sequence_length = max([len(seq) for seq in sequences]) padded_sequences = pad_sequences(sequences, maxlen=max_sequence_length, padding="post", truncating="post") return padded_sequences, np.array(labels) # Prepare training and test data X_train, y_train = prepare_sequences(train_df) X_test, y_test = prepare_sequences(test_df) # Check input shape (should be: [number_of_patients, max_sequence_length, 1]) print(f"Training input shape: {X_train.shape}") # Build the LSTM model model = Sequential() # Input shape is (timesteps, features) — sample count is inferred automatically model.add(LSTM(32, return_sequences=True, input_shape=(X_train.shape[1], X_train.shape[2]))) model.add(LSTM(64, return_sequences=False)) model.add(Dropout(0.5)) model.add(Dense(1)) model.add(Activation('sigmoid')) model.compile(loss='binary_crossentropy', optimizer='rmsprop') # Use a smaller batch size since we're now using patient-level samples model.fit(X_train, y_train, batch_size=32, epochs=11) score = model.evaluate(X_test, y_test, batch_size=32) print(f"Test loss: {score}")
Scenario 2: Treat each individual spectrum reading as a standalone sample (Not recommended for LSTMs)
If you really don't want to use sequential data and just want to predict labels from single spectrum values, LSTMs aren't the best tool (a simple dense network would be more efficient). But if you still need to use an LSTM, you can manually add a dummy timestep dimension to your input:
from keras.models import Sequential from keras.layers import Dense, Dropout, Activation, LSTM import numpy as np train = np.loadtxt("TrainDatasetFinal.txt", delimiter=",") test = np.loadtxt("testDatasetFinal.txt", delimiter=",") y_train = train[:,7] y_test = test[:,7] # Add a timestep dimension: shape becomes (1415684, 1, 1) train_spec = train[:,6].reshape(-1, 1, 1) test_spec = test[:,6].reshape(-1, 1, 1) # Build the model model = Sequential() # Input shape is (1 timestep, 1 feature) model.add(LSTM(32, return_sequences=True, input_shape=(1, 1))) model.add(LSTM(64, return_sequences=False)) model.add(Dropout(0.5)) model.add(Dense(1)) model.add(Activation('sigmoid')) model.compile(loss='binary_crossentropy', optimizer='rmsprop') model.fit(train_spec, y_train, batch_size=2000, epochs=11) score = model.evaluate(test_spec, y_test, batch_size=2000) print(f"Test loss: {score}")
A quick note: Scenario 1 is far more meaningful for LSTMs, as it lets the model learn how a patient's spectrum changes over time—this is the core strength of recurrent networks. Scenario 2 essentially wastes the LSTM's sequential capabilities.
内容的提问来源于stack exchange,提问作者Hadeer El-Zayat

