You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

LSTM语音情感识别特征维度不匹配错误求助

语音情感识别LSTM实现的维度不匹配问题解决

问题描述

我正尝试基于LSTM实现语音情感识别,训练矩阵维度为60093×39,测试矩阵维度为76503×39,各自对应独立的标签矩阵。运行Matlab代码时触发错误:The training sequences are of feature dimension 60093 but the input layer expects sequences of feature dimension 39。

原代码如下:

% Load the dataset
load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_app.mat');
load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_test.mat');

num_classes = 7;


% Transpose the training and test matrices
mfcc_matrix_app = mfcc_matrix_app';  % Transpose the training matrix
mfcc_matrix_test = mfcc_matrix_test';  % Transpose the test matrix

%featureSequences = HelperFeatureVector2Sequence(mfcc_matrix_app,20,5);
% Update the num_features variable with the correct value
num_features = size(mfcc_matrix_app, 2);
% Define the LSTM network architecture
num_hidden_units = 64;
layers = [
    sequenceInputLayer(num_features)
    lstmLayer(num_hidden_units, 'OutputMode', 'last')
    fullyConnectedLayer(num_classes)
    softmaxLayer
    classificationLayer];


% Specify the training options
max_epochs = 20;
mini_batch_size = 128;
initial_learning_rate = 0.001;

options = trainingOptions('adam', ...
    'MaxEpochs', max_epochs, ...
    'MiniBatchSize', mini_batch_size, ...
    'InitialLearnRate', initial_learning_rate, ...
    'GradientThreshold', 1, ...
    'Shuffle', 'every-epoch', ...
    'Verbose', 1, ...
    'Plots', 'training-progress');
`
% Train the LSTM network
net = trainNetwork(mfcc_matrix_app, CA, layers, options);

% Evaluate the performance of the trained model on the test data
predicted_labels = classify(net, mfcc_matrix_test);
accuracy = sum(predicted_labels == CT) / 7;
confusion_matrix = confusionmat(CT, predicted_labels);
disp(['Accuracy: ' num2str(accuracy)]);
disp('Confusion Matrix:');
disp(confusion_matrix);

工作区截图:
工作区截图

错误原因

Matlab中trainNetwork对序列输入的维度格式要求为**[特征数, 时间步长, 样本数](单样本序列为[特征数, 时间步长]**)。你的原始训练矩阵是60093×39,代表60093个样本、每个样本含39个MFCC特征,但转置后变成39×60093,导致num_features被错误设为60093,输入层期望的特征维度与实际输入完全颠倒,触发维度不匹配错误。

解决方案

  1. 不要转置原始矩阵,而是将扁平化的特征矩阵转换为LSTM需要的时序序列格式——语音是时序数据,每个样本应为一段固定长度的时序序列,每个时刻对应39维MFCC特征。
  2. 启用你注释掉的HelperFeatureVector2Sequence函数,它的作用就是将样本特征转换为符合要求的序列结构(参数20为窗口大小、5为步长,可根据你的语音帧设置调整)。
  3. 修正准确率计算逻辑,原代码除以7是错误的,应除以测试样本总数。

修改后的代码示例

% Load the dataset
load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_app.mat');
load('C:\Users\hamza\Desktop\deep_learning_matrix\matrice_mfcc_test.mat');

num_classes = 7;

% 将样本特征转换为时序序列,窗口大小20,步长5
featureSequences_app = HelperFeatureVector2Sequence(mfcc_matrix_app, 20, 5);
featureSequences_test = HelperFeatureVector2Sequence(mfcc_matrix_test, 20, 5);

% 获取正确的特征数(39)
num_features = size(featureSequences_app, 1);

% 定义LSTM网络结构
num_hidden_units = 64;
layers = [
    sequenceInputLayer(num_features)
    lstmLayer(num_hidden_units, 'OutputMode', 'last')
    fullyConnectedLayer(num_classes)
    softmaxLayer
    classificationLayer];

% 设置训练参数
max_epochs = 20;
mini_batch_size = 128;
initial_learning_rate = 0.001;

options = trainingOptions('adam', ...
    'MaxEpochs', max_epochs, ...
    'MiniBatchSize', mini_batch_size, ...
    'InitialLearnRate', initial_learning_rate, ...
    'GradientThreshold', 1, ...
    'Shuffle', 'every-epoch', ...
    'Verbose', 1, ...
    'Plots', 'training-progress');

% 训练网络
net = trainNetwork(featureSequences_app, CA, layers, options);

% 测试模型并计算性能
predicted_labels = classify(net, featureSequences_test);
accuracy = sum(predicted_labels == CT) / length(CT);
confusion_matrix = confusionmat(CT, predicted_labels);
disp(['Accuracy: ' num2str(accuracy)]);
disp('Confusion Matrix:');
disp(confusion_matrix);

额外注意:需确保标签矩阵CA、CT的样本数与转换后的序列特征集样本数一致。

内容的提问来源于stack exchange,提问作者Hamza

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 16:05:40