为何TF-IDF结合KNN的文本分类准确率远高于BoW与GloVe?
Great question! Let's break down the key reasons TF-IDF is delivering far higher accuracy (95%) compared to BoW (62%) and GloVe (57%) for your concept classification task with 107 test questions:
1. TF-IDF's Weighting Plays to KNN's Core Strength
KNN relies entirely on calculating distance between samples (like cosine similarity or Euclidean distance) to make classifications. Unlike BoW, which just counts raw word occurrences, TF-IDF assigns higher weights to words that are:
- Frequent in a specific concept's questions (high TF, term frequency)
- Rare across all concepts (low IDF, inverse document frequency)
These are exactly the words that best distinguish one concept from another. For example, if a term only appears regularly in questions tagged c1, TF-IDF amplifies its importance—so KNN will more easily group similar c1 questions together based on this meaningful signal.
2. BoW's Raw Counts Waste Signal on Non-Distinctive Words
BoW treats all words equally based on how often they appear, regardless of their classification value. Even after removing stopwords, you might still have words that show up frequently across all 4 concepts (e.g., generic terms related to your problem domain). BoW gives these unhelpful words high weights, which muddles the distance calculations KNN uses to make decisions. TF-IDF solves this by penalizing overused, non-distinctive words, focusing only on terms that drive classification.
3. GloVe's Semantic Approach Doesn't Fit Your Task's Context
GloVe is a pre-trained word vector model that captures general semantic relationships between words, but it's a poor fit for your scenario for two main reasons:
- Small dataset size: Your total dataset is likely around 535 samples (since 107 is 20% of the total). Pre-trained GloVe vectors are trained on massive corpora, and they don't adapt well to small, niche datasets where your specific concept vocabulary might not be well-represented in the original training data.
- Lost word importance: To create a sentence vector from GloVe, you typically average all word vectors in the question. This dilutes the impact of key distinguishing words—even if a critical term for
c3is present, its weight gets averaged out with other less important words, weakening the signal KNN needs to classify accurately.
4. Preprocessing Synergizes Better with TF-IDF
Your preprocessing steps (tokenization, stopword removal, stemming) work hand-in-hand with TF-IDF:
- Stemming groups variant forms of the same word (e.g., "run" and "running") into a single token, which makes TF-IDF's frequency counts more consistent and meaningful.
- Stopword removal eliminates words that TF-IDF would already assign low weights to, cleaning up the feature space without losing valuable signal.
BoW benefits from these steps too, but without the intelligent weighting of TF-IDF, it still can't prioritize the most useful terms. GloVe, on the other hand, may not align with your stemming choices—pre-trained GloVe vectors use their own tokenization, so your stemmed words might not map correctly to the pre-trained vectors, introducing noise.
内容的提问来源于stack exchange,提问作者Tay

