如何用已提取的MFCC训练SVM?普什图语语音识别毕设求助
Hey Qaiser, let's tackle your SVM implementation for your Pashto speech recognition project. You've already got the MFCC extraction sorted, so now we'll focus on turning those 12x34 MFCC matrices into a format SVM can use, training the classifier, and running predictions on your test set.
First, let's break down the steps you need: converting 2D MFCC features to 1D vectors (since MATLAB's SVM expects each sample as a row), preparing your labels, training the model, and evaluating performance. Here's the full workflow with code:
Step 1: Preprocess MFCC Features
Each of your MFCC outputs is a 12x34 matrix—we need to flatten this into a single 408-dimensional vector per sample so SVM can process it. We'll use cell arrays to store all your MFCCs first (either from your existing extraction code or a batch process we'll cover next).
%% 1. Flatten 2D MFCC to 1D feature vectors % Assume you have: % - train_mfcc_cell: 234x1 cell array, each element is 12x34 double % - test_mfcc_cell: 78x1 cell array, each element is 12x34 double % - train_labels: 234x1 vector (0-25, matching your vocabulary) % - test_labels: 78x1 vector (0-25, matching test set vocabulary) % Process training set num_train = length(train_mfcc_cell); train_features = zeros(num_train, 12*34); % 12*34 = 408 features per sample for i = 1:num_train train_features(i, :) = reshape(train_mfcc_cell{i}, 1, []); end % Process test set num_test = length(test_mfcc_cell); test_features = zeros(num_test, 12*34); for i = 1:num_test test_features(i, :) = reshape(test_mfcc_cell{i}, 1, []); end
Step 2: Train the SVM Classifier
We'll use MATLAB's fitcsvm function. I've included standard best practices like feature standardization (critical for SVM) and a linear kernel (great for starters—you can switch to RBF if you need better performance later).
%% 2. Train SVM Model svm_model = fitcsvm(train_features, train_labels, ... 'KernelFunction', 'linear', ... % Start with linear kernel; try 'rbf' if needed 'Standardize', true, ... % Normalize features to improve SVM performance 'Verbose', 1); % Show training progress
Step 3: Run Predictions on Test Set
Once trained, use predict to get labels for your test set, then calculate accuracy and visualize results with a confusion matrix.
%% 3. Predict & Evaluate [test_predictions, ~] = predict(svm_model, test_features); % Calculate accuracy accuracy = sum(test_predictions == test_labels)/num_test * 100; fprintf('Test Set Accuracy: %.2f%%\n', accuracy); % Plot confusion matrix to see misclassifications figure; confusionchart(test_labels, test_predictions); title('SVM Classifier Confusion Matrix');
Bonus: Batch MFCC Extraction (To Replace Your Manual Code)
Right now you're extracting MFCCs one file at a time—this batch code will save you time by processing all your training/test audio files in one go. Just update the folder paths and label logic to match your file naming convention.
%% Batch MFCC Extraction (Training + Test Sets) clear all; close all; clc; % MFCC Parameters (matches your existing code) Tw = 25; Ts = 10; alpha = 0.97; M = 40; C = 12; L = 22; LF = 0; HF = 4000; T = 16000; % ---------------------- % Process Training Set % ---------------------- train_audio_dir = 'path/to/your/training/wavs'; % Update this path train_audio_files = dir(fullfile(train_audio_dir, '*.wav')); train_mfcc_cell = cell(length(train_audio_files), 1); train_labels = zeros(length(train_audio_files), 1); % Assuming 9 samples per vocabulary (234 total / 26 vocab = 9) samples_per_vocab_train = 9; for i = 1:length(train_audio_files) % Read audio audio_path = fullfile(train_audio_dir, train_audio_files(i).name); [Data, fs, ~] = wavread(audio_path); Data = imresize(Data(:,1), [T, 1]); % Resize to fixed length % Extract MFCC [MFCCs, ~, ~] = mfcc(Data, fs, Tw, Ts, alpha, @hamming, [LF HF], M, C, L); train_mfcc_cell{i} = MFCCs; % Assign label (0-25) based on sample index train_labels(i) = floor((i-1)/samples_per_vocab_train); end % ---------------------- % Process Test Set % ---------------------- test_audio_dir = 'path/to/your/test/wavs'; % Update this path test_audio_files = dir(fullfile(test_audio_dir, '*.wav')); test_mfcc_cell = cell(length(test_audio_files), 1); test_labels = zeros(length(test_audio_files), 1); % 3 samples per vocabulary (78 total / 26 vocab = 3) samples_per_vocab_test = 3; for i = 1:length(test_audio_files) audio_path = fullfile(test_audio_dir, test_audio_files(i).name); [Data, fs, ~] = wavread(audio_path); Data = imresize(Data(:,1), [T, 1]); [MFCCs, ~, ~] = mfcc(Data, fs, Tw, Ts, alpha, @hamming, [LF HF], M, C, L); test_mfcc_cell{i} = MFCCs; test_labels(i) = floor((i-1)/samples_per_vocab_test); end
Quick Tips for Better Performance
- Try Different Kernels: If linear kernel accuracy is low, switch to
'rbf'(radial basis function) and tune theSigmaparameter, or useoptimizeHyperparametersto auto-tune. - Improve Labeling: Instead of assigning labels by index, extract them directly from your filenames (e.g., if your file is named
vocab_5_sample_2.wav, parse the5as the label) to avoid errors. - Data Augmentation: Since your dataset is small, add noise, adjust pitch/speed, or time-stretch your audio samples to generate more training data—this will help the model generalize better.
内容的提问来源于stack exchange,提问作者Qaiser Khan

