You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于帧提取特征构建适配CNN-LSTM的情绪识别视频数据集?

Hey there! Let's work through how to adapt your pre-extracted frame features to your CNN-LSTM model, and address your concern about the model distinguishing different videos. Here's a step-by-step breakdown:

1. Dataset Construction: Organize Frame Features into Video-Level Samples

Your initial idea of merging all CSVs into one will lose critical video boundary information—so we need to treat each video as a single sample (a sequence of frame features + its emotion label). Here's how to implement this:

Step 1: Load and label each video's frame features

First, iterate through all your CSV files, extract the emotion label from the filename, and collect frame features per video:

import pandas as pd
import glob
import numpy as np
import tensorflow as tf

# Path to your feature CSVs
feature_dir = "path/to/your/csv_files/"
feature_files = glob.glob(f"{feature_dir}/*.csv")

# Map emotion codes from filenames to numeric labels (adjust to your actual encoding)
def get_emotion_label(filename):
    # Example: if filename is "video_123_0.csv" where 0=angry, 1=happy, etc.
    emotion_code = filename.split("_")[-1].replace(".csv", "")
    label_map = {"0":0, "1":1, "2":2, "3":3, "4":4, "5":5}  # Match your 6 emotions
    return label_map[emotion_code]

# Collect video sequences and labels
video_sequences = []
labels = []

# Set a fixed sequence length (match the 30 in your model's Reshape layer)
MAX_FRAMES = 30
FEATURE_DIM = 2048  # Replace with your actual frame feature dimension

for file in feature_files:
    # Load frame features (assuming CSV has no index column, just feature columns)
    frame_features = pd.read_csv(file).values
    
    # Align sequence length: pad short videos, truncate long ones
    num_frames = frame_features.shape[0]
    if num_frames < MAX_FRAMES:
        # Pad with zeros at the end of the sequence
        padded_features = np.pad(frame_features, ((0, MAX_FRAMES - num_frames), (0,0)), mode="constant")
    else:
        # Truncate to the first MAX_FRAMES frames
        padded_features = frame_features[:MAX_FRAMES, :]
    
    video_sequences.append(padded_features)
    labels.append(get_emotion_label(file))

# Convert to numpy arrays for model input
X = np.array(video_sequences)  # Shape: (num_videos, MAX_FRAMES, FEATURE_DIM)
y = np.array(labels)  # Shape: (num_videos,)

For large datasets, avoid loading everything into memory by using tf.data.Dataset:

batch_size = 16

dataset = tf.data.Dataset.from_tensor_slices((X, y))
dataset = dataset.shuffle(buffer_size=len(X)) \
                 .batch(batch_size) \
                 .prefetch(tf.data.AUTOTUNE)
2. Adjust Your CNN-LSTM Model for Pre-Extracted Features

Your original model was built to process raw video frames (spatial input like (frames, height, width, channels)), but now we're feeding pre-extracted frame feature vectors. We need to tweak the model to fit this input while retaining your desired CNN-LSTM structure:

Option 1: Adapt the 3D Convolutions for feature vectors

Treat each frame's feature vector as a 1x1 spatial "image" with FEATURE_DIM channels. Adjust the 3D convolution kernels to operate on the feature dimension instead of spatial dimensions:

input_shape = (MAX_FRAMES, 1, 1, FEATURE_DIM)

model = tf.keras.Sequential()
model.add(tf.keras.layers.Input(shape=input_shape))

# Adjust Conv3D kernels to (1,1,k) to operate on the feature dimension
model.add(tf.keras.layers.Conv3D(6, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-1'))
model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-1'))
model.add(tf.keras.layers.Conv3D(16, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-2'))
model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-2'))
model.add(tf.keras.layers.Conv3D(32, (1, 1, 5), (1, 1, 1), activation='relu', name='conv-3'))
model.add(tf.keras.layers.AveragePooling3D((1, 1, 2), strides=(1, 1, 2), name='avgpool-3'))
model.add(tf.keras.layers.Conv3D(64, (1, 1, 4), (1, 1, 1), activation='relu', name='conv-4'))

# Reshape to (MAX_FRAMES, 64) to feed into LSTM (matches your original Reshape)
model.add(tf.keras.layers.Reshape((MAX_FRAMES, 64), name='reshape'))
model.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True, name='lstm-1'))
model.add(tf.keras.layers.CuDNNLSTM(64, name='lstm-2'))
model.add(tf.keras.layers.Dense(6, activation=tf.nn.softmax, name='result'))

# Compile and train
model.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
model.fit(dataset, epochs=20)

Option 2: Simplify the model (skip 3D Conv if it's redundant)

Since you already have pre-extracted frame features, the 3D convolution layer may be redundant. You can directly feed the feature sequences into LSTM with a pre-processing dense layer:

input_shape = (MAX_FRAMES, FEATURE_DIM)

model_simple = tf.keras.Sequential()
model_simple.add(tf.keras.layers.Input(shape=input_shape))
model_simple.add(tf.keras.layers.Dense(64, activation='relu'))  # Project features to 64-dim
model_simple.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True, name='lstm-1'))
model_simple.add(tf.keras.layers.CuDNNLSTM(64, name='lstm-2'))
model_simple.add(tf.keras.layers.Dense(6, activation='softmax', name='result'))

model_simple.compile(optimizer='adam', loss='sparse_categorical_crossentropy', metrics=['accuracy'])
model_simple.fit(dataset, epochs=20)
3. Will the Model Distinguish Between Different Videos?

Absolutely! Each sample in your dataset corresponds to a full video's sequence of frame features, paired with its unique emotion label. The LSTM layers are designed to learn temporal patterns within each video sequence—so during training, the model will associate distinct sequence patterns with specific emotions, effectively distinguishing between different videos.

If you want to handle variable-length sequences without padding, you can add a Masking layer to ignore padded zeros in shorter sequences:

model_with_mask = tf.keras.Sequential()
model_with_mask.add(tf.keras.layers.Masking(mask_value=0.0, input_shape=(None, FEATURE_DIM)))
model_with_mask.add(tf.keras.layers.Dense(64, activation='relu'))
model_with_mask.add(tf.keras.layers.CuDNNLSTM(64, return_sequences=True))
model_with_mask.add(tf.keras.layers.CuDNNLSTM(64))
model_with_mask.add(tf.keras.layers.Dense(6, activation='softmax'))

内容的提问来源于stack exchange,提问作者Mark Mahill

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 08:42:55