You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用已提取的MFCC训练SVM?普什图语语音识别毕设求助

MATLAB SVM Training & Prediction for Your Pashto ASR FYP

Hey Qaiser, let's tackle your SVM implementation for your Pashto speech recognition project. You've already got the MFCC extraction sorted, so now we'll focus on turning those 12x34 MFCC matrices into a format SVM can use, training the classifier, and running predictions on your test set.

First, let's break down the steps you need: converting 2D MFCC features to 1D vectors (since MATLAB's SVM expects each sample as a row), preparing your labels, training the model, and evaluating performance. Here's the full workflow with code:

Step 1: Preprocess MFCC Features

Each of your MFCC outputs is a 12x34 matrix—we need to flatten this into a single 408-dimensional vector per sample so SVM can process it. We'll use cell arrays to store all your MFCCs first (either from your existing extraction code or a batch process we'll cover next).

%% 1. Flatten 2D MFCC to 1D feature vectors
% Assume you have:
% - train_mfcc_cell: 234x1 cell array, each element is 12x34 double
% - test_mfcc_cell: 78x1 cell array, each element is 12x34 double
% - train_labels: 234x1 vector (0-25, matching your vocabulary)
% - test_labels: 78x1 vector (0-25, matching test set vocabulary)

% Process training set
num_train = length(train_mfcc_cell);
train_features = zeros(num_train, 12*34); % 12*34 = 408 features per sample
for i = 1:num_train
    train_features(i, :) = reshape(train_mfcc_cell{i}, 1, []);
end

% Process test set
num_test = length(test_mfcc_cell);
test_features = zeros(num_test, 12*34);
for i = 1:num_test
    test_features(i, :) = reshape(test_mfcc_cell{i}, 1, []);
end

Step 2: Train the SVM Classifier

We'll use MATLAB's fitcsvm function. I've included standard best practices like feature standardization (critical for SVM) and a linear kernel (great for starters—you can switch to RBF if you need better performance later).

%% 2. Train SVM Model
svm_model = fitcsvm(train_features, train_labels, ...
    'KernelFunction', 'linear', ...  % Start with linear kernel; try 'rbf' if needed
    'Standardize', true, ...         % Normalize features to improve SVM performance
    'Verbose', 1);                   % Show training progress

Step 3: Run Predictions on Test Set

Once trained, use predict to get labels for your test set, then calculate accuracy and visualize results with a confusion matrix.

%% 3. Predict & Evaluate
[test_predictions, ~] = predict(svm_model, test_features);

% Calculate accuracy
accuracy = sum(test_predictions == test_labels)/num_test * 100;
fprintf('Test Set Accuracy: %.2f%%\n', accuracy);

% Plot confusion matrix to see misclassifications
figure;
confusionchart(test_labels, test_predictions);
title('SVM Classifier Confusion Matrix');

Bonus: Batch MFCC Extraction (To Replace Your Manual Code)

Right now you're extracting MFCCs one file at a time—this batch code will save you time by processing all your training/test audio files in one go. Just update the folder paths and label logic to match your file naming convention.

%% Batch MFCC Extraction (Training + Test Sets)
clear all; close all; clc;

% MFCC Parameters (matches your existing code)
Tw = 25; Ts = 10; alpha = 0.97;
M = 40; C = 12; L = 22; LF = 0; HF = 4000; T = 16000;

% ----------------------
% Process Training Set
% ----------------------
train_audio_dir = 'path/to/your/training/wavs'; % Update this path
train_audio_files = dir(fullfile(train_audio_dir, '*.wav'));

train_mfcc_cell = cell(length(train_audio_files), 1);
train_labels = zeros(length(train_audio_files), 1);

% Assuming 9 samples per vocabulary (234 total / 26 vocab = 9)
samples_per_vocab_train = 9;
for i = 1:length(train_audio_files)
    % Read audio
    audio_path = fullfile(train_audio_dir, train_audio_files(i).name);
    [Data, fs, ~] = wavread(audio_path);
    Data = imresize(Data(:,1), [T, 1]); % Resize to fixed length
    
    % Extract MFCC
    [MFCCs, ~, ~] = mfcc(Data, fs, Tw, Ts, alpha, @hamming, [LF HF], M, C, L);
    train_mfcc_cell{i} = MFCCs;
    
    % Assign label (0-25) based on sample index
    train_labels(i) = floor((i-1)/samples_per_vocab_train);
end

% ----------------------
% Process Test Set
% ----------------------
test_audio_dir = 'path/to/your/test/wavs'; % Update this path
test_audio_files = dir(fullfile(test_audio_dir, '*.wav'));

test_mfcc_cell = cell(length(test_audio_files), 1);
test_labels = zeros(length(test_audio_files), 1);

% 3 samples per vocabulary (78 total / 26 vocab = 3)
samples_per_vocab_test = 3;
for i = 1:length(test_audio_files)
    audio_path = fullfile(test_audio_dir, test_audio_files(i).name);
    [Data, fs, ~] = wavread(audio_path);
    Data = imresize(Data(:,1), [T, 1]);
    
    [MFCCs, ~, ~] = mfcc(Data, fs, Tw, Ts, alpha, @hamming, [LF HF], M, C, L);
    test_mfcc_cell{i} = MFCCs;
    
    test_labels(i) = floor((i-1)/samples_per_vocab_test);
end

Quick Tips for Better Performance

  • Try Different Kernels: If linear kernel accuracy is low, switch to 'rbf' (radial basis function) and tune the Sigma parameter, or use optimizeHyperparameters to auto-tune.
  • Improve Labeling: Instead of assigning labels by index, extract them directly from your filenames (e.g., if your file is named vocab_5_sample_2.wav, parse the 5 as the label) to avoid errors.
  • Data Augmentation: Since your dataset is small, add noise, adjust pitch/speed, or time-stretch your audio samples to generate more training data—this will help the model generalize better.

内容的提问来源于stack exchange,提问作者Qaiser Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:54:09