Keras/TensorFlow中LSTM输入数据高效处理:无需手动滑窗方案问询
Great question—manually looping and appending sequences for LSTMs is not just tedious, it's also incredibly inefficient, especially with large datasets. You're absolutely right that there are better ways to handle this, both with vectorized operations and built-in Keras tools that eliminate the need for manual reshaping entirely.
1. Ditch Manual Loops for Vectorized Sliding Windows (NumPy/Pandas)
Your current loop uses pd.DataFrame.append, which is notoriously slow because it creates a new DataFrame every iteration. Instead, use NumPy's built-in sliding window functionality to generate all your sequences in one go—it's orders of magnitude faster and memory-efficient.
Here's how to adapt your code:
import numpy as np import pandas as pd # Convert your features DataFrame to a NumPy array (shape: [total_timesteps, num_features]) features_array = features.transpose().values numberOfTimesteps = 240 # Generate sliding window sequences in one vectorized operation # Output shape: [num_samples, numberOfTimesteps, num_features] slided_sequences = np.lib.stride_tricks.sliding_window_view( features_array, window_shape=(numberOfTimesteps,), axis=0 ) # If you still need a DataFrame (though LSTMs work directly with arrays), reshape it: lstmFeatures = pd.DataFrame(slided_sequences.reshape(-1, slided_sequences.shape[-1]))
This method uses memory strides instead of copying data, so it’s lightning-fast even for large datasets.
2. Let Keras Automatically Generate Sequences for You
If you want to skip precomputing all sequences entirely, Keras has a dedicated utility called tf.keras.utils.timeseries_dataset_from_array that dynamically creates LSTM-ready batches on the fly. This is perfect for large datasets where storing all sequences in memory isn't feasible.
Here's a complete example:
import tensorflow as tf # Convert features to a NumPy array (shape: [total_timesteps, num_features]) features_array = features.transpose().values numberOfTimesteps = 240 # Create a dataset of sequences ready for LSTM input # Adjust parameters based on your task (e.g., add targets if you're doing prediction) dataset = tf.keras.utils.timeseries_dataset_from_array( data=features_array, targets=None, # Set to features_array[numberOfTimesteps:] if predicting next timestep sequence_length=numberOfTimesteps, batch_size=32, # Choose a batch size that fits your memory shuffle=False # Keep sequential order for time-series tasks (set to True only if appropriate) ) # Define and train your LSTM directly on the dataset model = tf.keras.Sequential([ tf.keras.layers.LSTM(64, input_shape=(numberOfTimesteps, features_array.shape[-1])), tf.keras.layers.Dense(1) # Adjust output layer based on your task (regression/classification) ]) model.compile(optimizer='adam', loss='mean_squared_error') model.fit(dataset, epochs=10)
This tool handles all the reshaping behind the scenes—you just specify your sequence length, and it generates batches of shape [batch_size, timesteps, features] exactly as your LSTM expects.
Why Your Original Loop Was Slow
The pd.DataFrame.append method creates a new DataFrame every time you call it, which leads to exponential increases in memory usage and computation time as your dataset grows. Vectorized operations and Keras' dataset generator avoid this overhead entirely by working with the original data directly or generating batches dynamically.
内容的提问来源于stack exchange,提问作者JonathanSchmied

