LSTM语音情感识别特征维度不匹配错误求助
语音情感识别LSTM实现的维度不匹配问题解决
问题描述
我正尝试基于LSTM实现语音情感识别,训练矩阵维度为60093×39,测试矩阵维度为76503×39,各自对应独立的标签矩阵。运行Matlab代码时触发错误:The training sequences are of feature dimension 60093 but the input layer expects sequences of feature dimension 39。
原代码如下:
% Load the dataset load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_app.mat'); load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_test.mat'); num_classes = 7; % Transpose the training and test matrices mfcc_matrix_app = mfcc_matrix_app'; % Transpose the training matrix mfcc_matrix_test = mfcc_matrix_test'; % Transpose the test matrix %featureSequences = HelperFeatureVector2Sequence(mfcc_matrix_app,20,5); % Update the num_features variable with the correct value num_features = size(mfcc_matrix_app, 2); % Define the LSTM network architecture num_hidden_units = 64; layers = [ sequenceInputLayer(num_features) lstmLayer(num_hidden_units, 'OutputMode', 'last') fullyConnectedLayer(num_classes) softmaxLayer classificationLayer]; % Specify the training options max_epochs = 20; mini_batch_size = 128; initial_learning_rate = 0.001; options = trainingOptions('adam', ... 'MaxEpochs', max_epochs, ... 'MiniBatchSize', mini_batch_size, ... 'InitialLearnRate', initial_learning_rate, ... 'GradientThreshold', 1, ... 'Shuffle', 'every-epoch', ... 'Verbose', 1, ... 'Plots', 'training-progress'); ` % Train the LSTM network net = trainNetwork(mfcc_matrix_app, CA, layers, options); % Evaluate the performance of the trained model on the test data predicted_labels = classify(net, mfcc_matrix_test); accuracy = sum(predicted_labels == CT) / 7; confusion_matrix = confusionmat(CT, predicted_labels); disp(['Accuracy: ' num2str(accuracy)]); disp('Confusion Matrix:'); disp(confusion_matrix);
工作区截图:
错误原因
Matlab中trainNetwork对序列输入的维度格式要求为**[特征数, 时间步长, 样本数](单样本序列为[特征数, 时间步长]**)。你的原始训练矩阵是60093×39,代表60093个样本、每个样本含39个MFCC特征,但转置后变成39×60093,导致num_features被错误设为60093,输入层期望的特征维度与实际输入完全颠倒,触发维度不匹配错误。
解决方案
- 不要转置原始矩阵,而是将扁平化的特征矩阵转换为LSTM需要的时序序列格式——语音是时序数据,每个样本应为一段固定长度的时序序列,每个时刻对应39维MFCC特征。
- 启用你注释掉的
HelperFeatureVector2Sequence函数,它的作用就是将样本特征转换为符合要求的序列结构(参数20为窗口大小、5为步长,可根据你的语音帧设置调整)。 - 修正准确率计算逻辑,原代码除以7是错误的,应除以测试样本总数。
修改后的代码示例
% Load the dataset load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_app.mat'); load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_test.mat'); num_classes = 7; % 将样本特征转换为时序序列,窗口大小20,步长5 featureSequences_app = HelperFeatureVector2Sequence(mfcc_matrix_app, 20, 5); featureSequences_test = HelperFeatureVector2Sequence(mfcc_matrix_test, 20, 5); % 获取正确的特征数(39) num_features = size(featureSequences_app, 1); % 定义LSTM网络结构 num_hidden_units = 64; layers = [ sequenceInputLayer(num_features) lstmLayer(num_hidden_units, 'OutputMode', 'last') fullyConnectedLayer(num_classes) softmaxLayer classificationLayer]; % 设置训练参数 max_epochs = 20; mini_batch_size = 128; initial_learning_rate = 0.001; options = trainingOptions('adam', ... 'MaxEpochs', max_epochs, ... 'MiniBatchSize', mini_batch_size, ... 'InitialLearnRate', initial_learning_rate, ... 'GradientThreshold', 1, ... 'Shuffle', 'every-epoch', ... 'Verbose', 1, ... 'Plots', 'training-progress'); % 训练网络 net = trainNetwork(featureSequences_app, CA, layers, options); % 测试模型并计算性能 predicted_labels = classify(net, featureSequences_test); accuracy = sum(predicted_labels == CT) / length(CT); confusion_matrix = confusionmat(CT, predicted_labels); disp(['Accuracy: ' num2str(accuracy)]); disp('Confusion Matrix:'); disp(confusion_matrix);
额外注意:需确保标签矩阵CA、CT的样本数与转换后的序列特征集样本数一致。
内容的提问来源于stack exchange,提问作者Hamza
相关产品推荐
相关产品推荐

