如何用Scikit-learn为20 NewsGroups数据集构建词袋表示适配朴素贝叶斯分类器
Great choice working with the 20 NewsGroups dataset and Naive Bayes—it’s a classic combo for learning text classification! Let’s walk through how to implement Bag of Words (BoW) properly, then level up your model’s accuracy with some proven tweaks.
Step 1: Implement Bag of Words with CountVectorizer
The CountVectorizer from sklearn.feature_extraction.text is your go-to tool for BoW. It converts text into numerical feature vectors by counting word occurrences, and handles common preprocessing like lowercasing and stopword removal out of the box.
Here’s how to use it with your existing Bunch train/test data:
from sklearn.feature_extraction.text import CountVectorizer from sklearn.naive_bayes import MultinomialNB from sklearn.metrics import accuracy_score # Initialize vectorizer with stopword removal (filters out uninformative words like "the") vectorizer = CountVectorizer(stop_words='english', max_features=10000) # Fit on training data (learns vocabulary) and transform both train/test sets X_train_bow = vectorizer.fit_transform(train.data) X_test_bow = vectorizer.transform(test.data) # Train Multinomial Naive Bayes (ideal for count-based text data) clf = MultinomialNB() clf.fit(X_train_bow, train.target) # Evaluate accuracy y_pred = clf.predict(X_test_bow) print(f"BoW + MultinomialNB Accuracy: {accuracy_score(test.target, y_pred):.4f}")
Step 2: Upgrade to TF-IDF for Better Feature Representation
Raw word counts can overemphasize frequent but unimportant words. TF-IDF (Term Frequency-Inverse Document Frequency) adjusts weights to prioritize words that are unique to specific classes. You can use TfidfVectorizer to combine BoW and TF-IDF in one step:
from sklearn.feature_extraction.text import TfidfVectorizer # Initialize TF-IDF vectorizer with n-grams (captures phrases like "space shuttle") tfidf_vectorizer = TfidfVectorizer( stop_words='english', max_features=10000, ngram_range=(1, 2) # includes unigrams and bigrams ) X_train_tfidf = tfidf_vectorizer.fit_transform(train.data) X_test_tfidf = tfidf_vectorizer.transform(test.data) # Retrain and evaluate clf_tfidf = MultinomialNB() clf_tfidf.fit(X_train_tfidf, train.target) y_pred_tfidf = clf_tfidf.predict(X_test_tfidf) print(f"TF-IDF + MultinomialNB Accuracy: {accuracy_score(test.target, y_pred_tfidf):.4f}")
Step 3: Boost Accuracy with These Tweaks
Here are actionable ways to improve your model’s performance:
- Filter low-frequency words: Use
min_df=5in the vectorizer to ignore words that appear in fewer than 5 documents (reduces noise). - Tune Naive Bayes hyperparameters: The
alphaparameter controls smoothing—useGridSearchCVto find the optimal value:from sklearn.model_selection import GridSearchCV param_grid = {'alpha': [0.1, 0.5, 1.0, 2.0]} grid_search = GridSearchCV(MultinomialNB(), param_grid, cv=5) grid_search.fit(X_train_tfidf, train.target) print(f"Best alpha value: {grid_search.best_params_['alpha']}") print(f"Cross-validation accuracy: {grid_search.best_score_:.4f}") - Preprocess text further: Add steps like removing numbers, stemming (reducing words to their root form), or lemmatization using libraries like NLTK before vectorization.
- Try Bernoulli Naive Bayes: For binary feature representations (word present/absent instead of counts),
BernoulliNBmight yield better results for some text tasks.
Understanding Naive Bayes Workflow
To deepen your grasp, you can inspect the model’s internal logic:
- Use
clf.feature_log_prob_to see the log probability of each word belonging to a class. - Map word indices back to actual words with
vectorizer.get_feature_names_out()to identify which terms drive predictions for each newsgroup.
内容的提问来源于stack exchange,提问作者tushariyer

