基于SVM的多分类任务:准确率与F值计算及评估合理性咨询
Great questions—let’s break this down step by step, starting with the general multi-class case then applying it directly to your SVM scenario.
First, let’s cover the core metrics when dealing with categorical variables that have more than two classes:
Accuracy
Accuracy is straightforward even for multi-class tasks: it’s the proportion of total samples that were predicted correctly. The formula is:
Accuracy = (Sum of True Positives for all classes) / Total Number of Samples
For example, if you have 3 classes and correctly predict 80 out of 100 samples, your accuracy is 80%. Just note: accuracy can be misleading if your dataset is imbalanced (one class has way more samples than others)—a model could just predict the majority class and get high accuracy without actually learning to distinguish classes.
F-Score (F1-Score)
Since F-score is the harmonic mean of precision and recall, for multi-class tasks we have two standard ways to compute it:
Macro-averaged F1: Treats every class equally, regardless of how many samples it has.
- Calculate precision and recall for each individual class.
- Take the arithmetic mean of all class precisions, and the arithmetic mean of all class recalls.
- Compute the harmonic mean of these two averages to get the macro F1-score.
This is ideal when every class is equally important to your use case.
Micro-averaged F1: Gives more weight to classes with more samples.
- Sum up the true positives (TP), false positives (FP), and false negatives (FN) across all classes.
- Use these totals to compute a global precision and global recall.
- Calculate the harmonic mean of these global values for the micro F1-score.
This works best if you care more about overall prediction correctness than individual class performance, especially with imbalanced datasets.
Your task is using a 4-class input (Building Type: Residential, Commercial, Industry, Special Buildings) to predict a 3-class output (Population Classification: High, MED, LOW). Here’s how to compute the metrics:
Accuracy Calculation
- Count all samples where the model’s prediction matches the actual population class:
- Number of samples where predicted = High AND actual = High
- Plus number of samples where predicted = MED AND actual = MED
- Plus number of samples where predicted = LOW AND actual = LOW
- Divide this total by the number of samples you’re evaluating (e.g., your test set size).
F-Score Calculation
Choose between macro or micro F1 based on your priorities:
- If High, MED, LOW are all equally important: Use macro F1. For each population class, calculate its precision (how many predicted High samples were actually High) and recall (how many actual High samples were correctly predicted), compute each class’s F1, then average those three F1 scores.
- If your dataset is imbalanced (e.g., High has 70% of samples, LOW has 10%): Use micro F1. Sum TP, FP, FN across all three population classes, then compute global precision/recall and their harmonic mean.
Is Accuracy Suitable for Multi-Class Evaluation?
Yes, absolutely—but it’s not always sufficient. Accuracy works well when your classes are roughly balanced, but it can be deceptive if one class dominates the dataset. For example, if 90% of your samples are MED, a model that just predicts MED for everything will have 90% accuracy, but it’s useless for identifying High or LOW. Always pair accuracy with F-scores, confusion matrices, or class-specific precision/recall to get a full picture of your model’s performance.
内容的提问来源于stack exchange,提问作者vaibhav kumar

