You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Multinomial Naïve Bayes的Python新闻标题文本分类实现问询

Got it, let's walk through building this Multinomial Naïve Bayes text classifier step by step in Jupyter Notebook. I'll break this down into actionable parts so you can follow along easily:

1. Prepare Your Dataset (4 News Categories)

First, you need to build your dataset of news titles and save them into category-named text files:

  • Create four plain text files: Politics.txt, Sports.txt, Weather.txt, Business.txt
  • Populate each file with at least 100 unique news titles, one title per line. You can source these from news archives, public datasets, or even generate realistic titles manually.
  • Example snippet for Sports.txt:

Golden State Warriors Clinch Playoff Spot with Dominant Road Win
Serena Williams Announces Comeback for 2024 Grand Slam Tournaments
Formula 1 Driver Max Verstappen Takes Pole Position for Monaco Grand Prix

2. Implement the Multinomial Naïve Bayes Classifier

We'll use scikit-learn's built-in tools to handle feature extraction and model training. Here's the complete code:

Step 1: Import Required Libraries

import numpy as np
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score

Step 2: Load and Preprocess the Dataset

# Map categories to their respective file paths
categories = ['Politics', 'Sports', 'Weather', 'Business']
file_paths = ['Politics.txt', 'Sports.txt', 'Weather.txt', 'Business.txt']

# Initialize lists to store text data and labels
texts = []
labels = []

# Load data from each file
for category, path in zip(categories, file_paths):
    with open(path, 'r', encoding='utf-8') as file:
        for line in file:
            cleaned_title = line.strip()
            if cleaned_title:  # Skip empty lines to avoid invalid samples
                texts.append(cleaned_title)
                labels.append(category)

Step 3: Convert Text to Numerical Features

We'll use CountVectorizer to convert raw text into word frequency vectors, and remove common English stopwords (like "the", "and") to reduce noise:

vectorizer = CountVectorizer(stop_words='english')
X = vectorizer.fit_transform(texts)  # Features: word frequency matrix
y = np.array(labels)  # Labels: news categories

Step 4: Split Data and Train the Model

# Split dataset into training (80%) and testing (20%) sets
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

# Initialize and train the Multinomial Naïve Bayes model
nb_model = MultinomialNB()
nb_model.fit(X_train, y_train)

# Evaluate model accuracy on test data
y_pred = nb_model.predict(X_test)
print(f"Model Accuracy on Test Set: {accuracy_score(y_test, y_pred):.2f}")

Step 5: Build a User Input Predictor

Add a function to take user input and return the predicted category:

def predict_news_category(news_title):
    # Convert input title to feature vector using the trained vectorizer
    title_vector = vectorizer.transform([news_title])
    # Predict and return the category
    predicted_category = nb_model.predict(title_vector)[0]
    return predicted_category

# Example usage
user_input = input("Enter a news title to classify: ")
result = predict_news_category(user_input)
print(f"This news title belongs to the **{result}** category.")
3. Pro Tips for Better Performance
  • Improve Feature Extraction: Replace CountVectorizer with TfidfVectorizer to prioritize important words over frequent but low-information words.
  • Handle Imbalanced Data: If one category has way more titles than others, add class_weight='balanced' to MultinomialNB() to adjust for this.
  • Clean Text: Add extra preprocessing steps like lowercasing, removing special characters, or stemming to standardize your input.

内容的提问来源于stack exchange,提问作者Pooja Khatri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:14:11