You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scikit-learn为20 NewsGroups数据集构建词袋表示适配朴素贝叶斯分类器

Guide to Bag of Words & Naive Bayes for 20 NewsGroups in Scikit-learn

Great choice working with the 20 NewsGroups dataset and Naive Bayes—it’s a classic combo for learning text classification! Let’s walk through how to implement Bag of Words (BoW) properly, then level up your model’s accuracy with some proven tweaks.

Step 1: Implement Bag of Words with CountVectorizer

The CountVectorizer from sklearn.feature_extraction.text is your go-to tool for BoW. It converts text into numerical feature vectors by counting word occurrences, and handles common preprocessing like lowercasing and stopword removal out of the box.

Here’s how to use it with your existing Bunch train/test data:

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import accuracy_score

# Initialize vectorizer with stopword removal (filters out uninformative words like "the")
vectorizer = CountVectorizer(stop_words='english', max_features=10000)

# Fit on training data (learns vocabulary) and transform both train/test sets
X_train_bow = vectorizer.fit_transform(train.data)
X_test_bow = vectorizer.transform(test.data)

# Train Multinomial Naive Bayes (ideal for count-based text data)
clf = MultinomialNB()
clf.fit(X_train_bow, train.target)

# Evaluate accuracy
y_pred = clf.predict(X_test_bow)
print(f"BoW + MultinomialNB Accuracy: {accuracy_score(test.target, y_pred):.4f}")

Step 2: Upgrade to TF-IDF for Better Feature Representation

Raw word counts can overemphasize frequent but unimportant words. TF-IDF (Term Frequency-Inverse Document Frequency) adjusts weights to prioritize words that are unique to specific classes. You can use TfidfVectorizer to combine BoW and TF-IDF in one step:

from sklearn.feature_extraction.text import TfidfVectorizer

# Initialize TF-IDF vectorizer with n-grams (captures phrases like "space shuttle")
tfidf_vectorizer = TfidfVectorizer(
    stop_words='english',
    max_features=10000,
    ngram_range=(1, 2)  # includes unigrams and bigrams
)

X_train_tfidf = tfidf_vectorizer.fit_transform(train.data)
X_test_tfidf = tfidf_vectorizer.transform(test.data)

# Retrain and evaluate
clf_tfidf = MultinomialNB()
clf_tfidf.fit(X_train_tfidf, train.target)
y_pred_tfidf = clf_tfidf.predict(X_test_tfidf)
print(f"TF-IDF + MultinomialNB Accuracy: {accuracy_score(test.target, y_pred_tfidf):.4f}")

Step 3: Boost Accuracy with These Tweaks

Here are actionable ways to improve your model’s performance:

  • Filter low-frequency words: Use min_df=5 in the vectorizer to ignore words that appear in fewer than 5 documents (reduces noise).
  • Tune Naive Bayes hyperparameters: The alpha parameter controls smoothing—use GridSearchCV to find the optimal value:
    from sklearn.model_selection import GridSearchCV
    
    param_grid = {'alpha': [0.1, 0.5, 1.0, 2.0]}
    grid_search = GridSearchCV(MultinomialNB(), param_grid, cv=5)
    grid_search.fit(X_train_tfidf, train.target)
    
    print(f"Best alpha value: {grid_search.best_params_['alpha']}")
    print(f"Cross-validation accuracy: {grid_search.best_score_:.4f}")
    
  • Preprocess text further: Add steps like removing numbers, stemming (reducing words to their root form), or lemmatization using libraries like NLTK before vectorization.
  • Try Bernoulli Naive Bayes: For binary feature representations (word present/absent instead of counts), BernoulliNB might yield better results for some text tasks.

Understanding Naive Bayes Workflow

To deepen your grasp, you can inspect the model’s internal logic:

  • Use clf.feature_log_prob_ to see the log probability of each word belonging to a class.
  • Map word indices back to actual words with vectorizer.get_feature_names_out() to identify which terms drive predictions for each newsgroup.

内容的提问来源于stack exchange,提问作者tushariyer

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:19:56