基于Matlab fitcdiscr的单特征LDA分类及边界可视化问询
Great question! Let's break this down step by step for your single-feature LDA classification task in MATLAB:
For a single-feature scenario, the most intuitive visualization combines the distribution of both classes with the LDA decision boundary. This lets you clearly see how well the feature separates the two groups and where the classifier draws its line.
Here’s a complete code example with explanations:
% Step 1: Simulate sample data (replace with your own dataset) rng(1); % Fix random seed for reproducibility n_class1 = 50; n_class2 = 50; X_class1 = normrnd(2, 1, n_class1, 1); % Feature values for class 1 X_class2 = normrnd(5, 1, n_class2, 1); % Feature values for class 2 X = [X_class1; X_class2]; % Combined feature matrix (n_samples x 1) Y = [ones(n_class1,1); 2*ones(n_class2,1)]; % Class labels (1 and 2) % Step 2: Train LDA classifier with a single feature lda_model = fitcdiscr(X, Y); % Step 3: Visualize class distributions and decision boundary figure('Position', [100 100 800 500]); hold on; box on; % Plot kernel density estimates for both classes (smoother than histograms) [f1, x_vals1] = ksdensity(X_class1); [f2, x_vals2] = ksdensity(X_class2); plot(x_vals1, f1, 'b', 'LineWidth', 2); plot(x_vals2, f2, 'r', 'LineWidth', 2); % Calculate and plot the decision boundary threshold coeffs = lda_model.Coefficients; decision_threshold = -coeffs.Intercept / coeffs.Linear; plot([decision_threshold decision_threshold], [0 max([max(f1), max(f2)])], 'k--', 'LineWidth', 2); % Add labels and legend xlabel('Feature Value', 'FontSize', 12); ylabel('Density', 'FontSize', 12); title('Single-Feature LDA Classification Performance', 'FontSize', 14); legend('Class 1', 'Class 2', 'Decision Boundary', 'Location', 'best'); hold off;
This plot shows you how much overlap exists between the two classes and whether the boundary is placed logically to separate them.
fitcdiscr For single-feature LDA, the decision boundary is a single threshold value (a vertical line on your plot). You can extract this directly from the trained model:
- The LDA decision function takes the form
g(x) = coeffs.Linear * x + coeffs.Intercept - The boundary occurs where
g(x) = 0, so solving for x givesx = -coeffs.Intercept / coeffs.Linear
If you ever work with multiple features, the boundary becomes a hyperplane, but for your single-feature use case, this simple threshold is all you need. You can also verify it by generating a range of x values, predicting their classes with predict(lda_model, x_range), and finding where the class prediction switches.
fitcdiscr You do NOT need to explicitly input the number of samples! The fitcdiscr function automatically infers sample count from the dimensions of your input data:
X: A matrix where each row is a sample, each column is a feature (for your task, this will be ann_samples x 1matrix)Y: A vector of class labels with exactlyn_sampleselements (one label per row inX)
Just call lda_model = fitcdiscr(X, Y) and the function handles the rest.
Pro tip: To quantify each feature's classification performance (not just visualize it), use cross-validation:
cv_model = fitcdiscr(X, Y, 'CrossVal', 'On'); classification_error = kfoldLoss(cv_model);
This gives you a numerical score to compare features and pick the optimal one.
内容的提问来源于stack exchange,提问作者Feng Ma

