在R语言中实现两类朴素贝叶斯文本分类遇到问题
Hey Peter, sounds like you've got a solid start with your text preprocessing using the tm package—great job on that! Let's walk through how to implement a Naive Bayes classifier for your dataset (1k samples per class for testing, scalable to the full 250k later) step by step.
1. Prepare Your Preprocessed Text Data
First, make sure your preprocessed text is converted into a model-ready format. If you're using tm, you probably have a DocumentTermMatrix (DTM) object. Let's convert that to a data frame and attach your classification labels:
# Load required packages library(tm) library(e1071) # Go-to package for Naive Bayes in R # Convert DTM to a data frame (replace `dtm` with your actual DTM object) dtm_df <- as.data.frame(as.matrix(dtm)) # Add your classification label (replace `label` with your actual label variable; values: "Complaint"/"General_Text") dtm_df$label <- factor(label, levels = c("Complaint", "General_Text"))
2. Split Data into Training & Test Sets
Since you're working with 1k samples per class, let's split into a 70/30 train/test split to evaluate model performance:
set.seed(123) # Set seed for reproducible results train_idx <- sample(1:nrow(dtm_df), 0.7 * nrow(dtm_df)) train_data <- dtm_df[train_idx, ] test_data <- dtm_df[-train_idx, ] # Verify class balance in splits cat("Training set class distribution:\n") print(table(train_data$label)) cat("\nTest set class distribution:\n") print(table(test_data$label))
3. Train the Naive Bayes Model
The naiveBayes() function from e1071 handles text data well, including built-in Laplace smoothing to avoid zero-probability issues with rare words:
# Train the model (predict `label` using all other columns as features) nb_model <- naiveBayes(label ~ ., data = train_data) # Optional: Adjust Laplace smoothing if needed (default is 0; try 1 for stronger smoothing) # nb_model <- naiveBayes(label ~ ., data = train_data, laplace = 1)
4. Evaluate Model Performance
Now let's test the model on the test set and calculate key metrics:
# Generate predictions predictions <- predict(nb_model, newdata = test_data) # Create a confusion matrix confusion_mat <- table(Actual = test_data$label, Predicted = predictions) cat("Confusion Matrix:\n") print(confusion_mat) # Calculate accuracy accuracy <- sum(diag(confusion_mat)) / sum(confusion_mat) cat("\nModel Accuracy:", round(accuracy * 100, 2), "%\n")
5. Scaling to Full 250k Samples per Class
When you move to the full dataset, you might run into memory issues with a large DTM. Here are two fixes:
- Remove sparse terms: Use
removeSparseTerms()to drop words that appear in very few documents (e.g., keep terms that appear in at least 1% of documents):dtm_sparse <- removeSparseTerms(dtm, 0.99) # 0.99 = keep terms with sparsity ≤ 99% - Use a more efficient package: For large-scale text data,
text2vecis faster and more memory-efficient thantm. It works seamlessly with Naive Bayes implementations too.
Quick Troubleshooting Tips
If your model performance is underwhelming:
- Double-check your preprocessing: Randomly sample 5-10 preprocessed texts to confirm stopwords, punctuation, numbers, and irrelevant terms are removed, and stemming/lemmatization is applied correctly.
- Adjust smoothing: If the model is overfitting (great on training, bad on test), increase the
laplaceparameter. If it's underfitting, decrease it. - Verify class balance: Even if your full dataset is balanced, ensure your train/test splits don't accidentally skew class ratios.
内容的提问来源于stack exchange,提问作者Peter Holler

