基于Multinomial Naïve Bayes的Python新闻标题文本分类实现问询
Got it, let's walk through building this Multinomial Naïve Bayes text classifier step by step in Jupyter Notebook. I'll break this down into actionable parts so you can follow along easily:
First, you need to build your dataset of news titles and save them into category-named text files:
- Create four plain text files:
Politics.txt,Sports.txt,Weather.txt,Business.txt - Populate each file with at least 100 unique news titles, one title per line. You can source these from news archives, public datasets, or even generate realistic titles manually.
- Example snippet for
Sports.txt:
Golden State Warriors Clinch Playoff Spot with Dominant Road Win
Serena Williams Announces Comeback for 2024 Grand Slam Tournaments
Formula 1 Driver Max Verstappen Takes Pole Position for Monaco Grand Prix
We'll use scikit-learn's built-in tools to handle feature extraction and model training. Here's the complete code:
Step 1: Import Required Libraries
import numpy as np from sklearn.feature_extraction.text import CountVectorizer from sklearn.naive_bayes import MultinomialNB from sklearn.model_selection import train_test_split from sklearn.metrics import accuracy_score
Step 2: Load and Preprocess the Dataset
# Map categories to their respective file paths categories = ['Politics', 'Sports', 'Weather', 'Business'] file_paths = ['Politics.txt', 'Sports.txt', 'Weather.txt', 'Business.txt'] # Initialize lists to store text data and labels texts = [] labels = [] # Load data from each file for category, path in zip(categories, file_paths): with open(path, 'r', encoding='utf-8') as file: for line in file: cleaned_title = line.strip() if cleaned_title: # Skip empty lines to avoid invalid samples texts.append(cleaned_title) labels.append(category)
Step 3: Convert Text to Numerical Features
We'll use CountVectorizer to convert raw text into word frequency vectors, and remove common English stopwords (like "the", "and") to reduce noise:
vectorizer = CountVectorizer(stop_words='english') X = vectorizer.fit_transform(texts) # Features: word frequency matrix y = np.array(labels) # Labels: news categories
Step 4: Split Data and Train the Model
# Split dataset into training (80%) and testing (20%) sets X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # Initialize and train the Multinomial Naïve Bayes model nb_model = MultinomialNB() nb_model.fit(X_train, y_train) # Evaluate model accuracy on test data y_pred = nb_model.predict(X_test) print(f"Model Accuracy on Test Set: {accuracy_score(y_test, y_pred):.2f}")
Step 5: Build a User Input Predictor
Add a function to take user input and return the predicted category:
def predict_news_category(news_title): # Convert input title to feature vector using the trained vectorizer title_vector = vectorizer.transform([news_title]) # Predict and return the category predicted_category = nb_model.predict(title_vector)[0] return predicted_category # Example usage user_input = input("Enter a news title to classify: ") result = predict_news_category(user_input) print(f"This news title belongs to the **{result}** category.")
- Improve Feature Extraction: Replace
CountVectorizerwithTfidfVectorizerto prioritize important words over frequent but low-information words. - Handle Imbalanced Data: If one category has way more titles than others, add
class_weight='balanced'toMultinomialNB()to adjust for this. - Clean Text: Add extra preprocessing steps like lowercasing, removing special characters, or stemming to standardize your input.
内容的提问来源于stack exchange,提问作者Pooja Khatri

