如何基于帧提取特征构建适配CNN-LSTM的情绪识别视频数据集?
Hey there! Let's work through how to adapt your pre-extracted frame features to your CNN-LSTM model, and address your concern about the model distinguishing different videos. Here's a step-by-step breakdown:
Your initial idea of merging all CSVs into one will lose critical video boundary information—so we need to treat each video as a single sample (a sequence of frame features + its emotion label). Here's how to implement this:
Step 1: Load and label each video's frame features
First, iterate through all your CSV files, extract the emotion label from the filename, and collect frame features per video:
import pandas as pd import glob import numpy as np import tensorflow as tf # Path to your feature CSVs feature_dir = "path/to/your/csv_files/" feature_files = glob.glob(f"{feature_dir}/*.csv") # Map emotion codes from filenames to numeric labels (adjust to your actual encoding) def get_emotion_label(filename): # Example: if filename is "video_123_0.csv" where 0=angry, 1=happy, etc. emotion_code = filename.split("_")[-1].replace(".csv", "") label_map = {"0":0, "1":1, "2":2, "3":3, "4":4, "5":5} # Match your 6 emotions return label_map[emotion_code] # Collect video sequences and labels video_sequences = [] labels = [] # Set a fixed sequence length (match the 30 in your model's Reshape layer) MAX_FRAMES = 30 FEATURE_DIM = 2048 # Replace with your actual frame feature dimension for file in feature_files: # Load frame features (assuming CSV has no index column, just feature columns) frame_features = pd.read_csv(file).values # Align sequence length: pad short videos, truncate long ones num_frames = frame_features.shape[0] if num_frames < MAX_FRAMES: # Pad with zeros at the end of the sequence padded_features = np.pad(frame_features, ((0, MAX_FRAMES - num_frames), (0,0)), mode="constant") else: # Truncate to the first MAX_FRAMES frames padded_features = frame_features[:MAX_FRAMES, :] video_sequences.append(padded_features) labels.append(get_emotion_label(file)) # Convert to numpy arrays for model input X = np.array(video_sequences) # Shape: (num_videos, MAX_FRAMES, FEATURE_DIM) y = np.array(labels) # Shape: (num_videos,)
Step 2: Use TensorFlow Dataset for efficient processing (optional but recommended)
For large datasets, avoid loading everything into memory by using tf.data.Dataset:
batch_size = 16 dataset = tf.data.Dataset.from_tensor_slices((X, y)) dataset = dataset.shuffle(buffer_size=len(X)) \ .batch(batch_size) \ .prefetch(tf.data.AUTOTUNE)
Your original model was built to process raw video frames (spatial input like (frames, height, width, channels)), but now we're feeding pre-extracted frame feature vectors. We need to tweak the model to fit this input while retaining your desired CNN-LSTM structure:
Option 1: Adapt the 3D Convolutions for feature vectors
Treat each frame's feature vector as a 1x1 spatial "image" with FEATURE_DIM channels. Adjust the 3D convolution kernels to operate on the feature dimension instead of spatial dimensions:
input_shape = (MAX_FRAMES, 1, 1, FEATURE_DIM) model = tf.keras.Sequential() model.add(tf.keras.layers.Input(shape=input_shape)) # Adjust Conv3D kernels to (1,1,k) to operate on the feature dimension model.add(tf.keras.layers.Conv3D(6, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-1')) model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-1')) model.add(tf.keras.layers.Conv3D(16, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-2')) model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-2')) model.add(tf.keras.layers.Conv3D(32, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-3')) model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-3')) model.add(tf.keras.layers.Conv3D(64, (1, 1, 4), (1, 1, 1), activation='relu', name='conv-4')) # Reshape to (MAX_FRAMES, 64) to feed into LSTM (matches your original Reshape) model.add(tf.keras.layers.Reshape((MAX_FRAMES, 64), name='reshape')) model.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True, name='lstm-1')) model.add(tf.keras.layers.CuDNNLSTM(64, name='lstm-2')) model.add(tf.keras.layers.Dense(6, activation=tf.nn.softmax, name='result')) # Compile and train model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) model.fit(dataset, epochs=20)
Option 2: Simplify the model (skip 3D Conv if it's redundant)
Since you already have pre-extracted frame features, the 3D convolution layer may be redundant. You can directly feed the feature sequences into LSTM with a pre-processing dense layer:
input_shape = (MAX_FRAMES, FEATURE_DIM) model_simple = tf.keras.Sequential() model_simple.add(tf.keras.layers.Input(shape=input_shape)) model_simple.add(tf.keras.layers.Dense(64, activation='relu')) # Project features to 64-dim model_simple.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True, name='lstm-1')) model_simple.add(tf.keras.layers.CuDNNLSTM(64, name='lstm-2')) model_simple.add(tf.keras.layers.Dense(6, activation='softmax', name='result')) model_simple.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy']) model_simple.fit(dataset, epochs=20)
Absolutely! Each sample in your dataset corresponds to a full video's sequence of frame features, paired with its unique emotion label. The LSTM layers are designed to learn temporal patterns within each video sequence—so during training, the model will associate distinct sequence patterns with specific emotions, effectively distinguishing between different videos.
If you want to handle variable-length sequences without padding, you can add a Masking layer to ignore padded zeros in shorter sequences:
model_with_mask = tf.keras.Sequential() model_with_mask.add(tf.keras.layers.Masking(mask_value=0.0, input_shape=(None, FEATURE_DIM))) model_with_mask.add(tf.keras.layers.Dense(64, activation='relu')) model_with_mask.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True)) model_with_mask.add(tf.keras.layers.CuDNNLSTM(64)) model_with_mask.add(tf.keras.layers.Dense(6, activation='softmax'))
内容的提问来源于stack exchange,提问作者Mark Mahill

