如何判断分类器属于高偏差(High Bias)还是高方差(High Variance)?
Alright, let’s tackle this question head-on—since understanding how to spot high bias vs. high variance is key to fixing underfitting and overfitting in classifiers. You already know the basics of the bias-variance tradeoff, so let’s jump straight to the practical telltale signs:
How to Spot High Bias (Underfitting)
High bias means your model is too simple to capture the underlying pattern in the data. Here’s what to look for:
- Both training and test performance are poor: Your model can’t even learn the basic relationships in the training data, let alone generalize to new data. For example, using a linear regression model to classify data with a nonlinear decision boundary will result in low accuracy across both sets.
- Model is overly simplistic: Think shallow decision trees, linear models with too few features, or models with heavy regularization cranked up too high.
- Adding more training data does nothing: The problem isn’t a lack of data—it’s that the model doesn’t have the capacity to learn the pattern. Throwing more data at it won’t close the performance gap.
How to Spot High Variance (Overfitting)
High variance means your model has memorized noise and random fluctuations in the training data instead of learning the true underlying pattern. Watch for these signs:
- Training performance is great, test performance is terrible: Your model nails the training set (even gets near-perfect accuracy) but falls apart when presented with unseen data. This is the classic "memorization vs. generalization" red flag.
- Model is overly complex: Think deep decision trees with no pruning, unregularized neural networks, or models using hundreds of high-dimensional features without any dimensionality reduction.
- Adding more training data improves performance: Since the model’s overfitting comes from limited data that doesn’t represent the true distribution, expanding the training set helps it learn the real pattern instead of noise.
Your Note About Insufficient Training Data
You mentioned cases where the data lacks enough information related to the target function (i.e., small sample size)—this directly amplifies variance. When you have too few samples, the model can easily be skewed by random quirks in the training data: for example, if your training set happens to have an unusually high proportion of one class’s edge cases, the model will treat those quirks as the "normal" pattern. When it encounters real-world data that follows the true distribution, it fails miserably. Even a moderately complex model can suffer from high variance if it’s trained on too little data.
Quick Practical Check: Learning Curves
A surefire way to confirm is to plot learning curves (training error vs. validation error as training sample size increases):
- If both curves are high and stay close together → High bias (underfitting)
- If training error is low, validation error is high, and the gap between them is large → High variance (overfitting)
内容的提问来源于stack exchange,提问作者Ébe Isaac

