You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R语言中实现两类朴素贝叶斯文本分类遇到问题

Hey Peter, sounds like you've got a solid start with your text preprocessing using the tm package—great job on that! Let's walk through how to implement a Naive Bayes classifier for your dataset (1k samples per class for testing, scalable to the full 250k later) step by step.

Implementing a Naive Bayes Classifier for Your Text Dataset

1. Prepare Your Preprocessed Text Data

First, make sure your preprocessed text is converted into a model-ready format. If you're using tm, you probably have a DocumentTermMatrix (DTM) object. Let's convert that to a data frame and attach your classification labels:

# Load required packages
library(tm)
library(e1071) # Go-to package for Naive Bayes in R

# Convert DTM to a data frame (replace `dtm` with your actual DTM object)
dtm_df <- as.data.frame(as.matrix(dtm))
# Add your classification label (replace `label` with your actual label variable; values: "Complaint"/"General_Text")
dtm_df$label <- factor(label, levels = c("Complaint", "General_Text"))

2. Split Data into Training & Test Sets

Since you're working with 1k samples per class, let's split into a 70/30 train/test split to evaluate model performance:

set.seed(123) # Set seed for reproducible results
train_idx <- sample(1:nrow(dtm_df), 0.7 * nrow(dtm_df))
train_data <- dtm_df[train_idx, ]
test_data <- dtm_df[-train_idx, ]

# Verify class balance in splits
cat("Training set class distribution:\n")
print(table(train_data$label))
cat("\nTest set class distribution:\n")
print(table(test_data$label))

3. Train the Naive Bayes Model

The naiveBayes() function from e1071 handles text data well, including built-in Laplace smoothing to avoid zero-probability issues with rare words:

# Train the model (predict `label` using all other columns as features)
nb_model <- naiveBayes(label ~ ., data = train_data)

# Optional: Adjust Laplace smoothing if needed (default is 0; try 1 for stronger smoothing)
# nb_model <- naiveBayes(label ~ ., data = train_data, laplace = 1)

4. Evaluate Model Performance

Now let's test the model on the test set and calculate key metrics:

# Generate predictions
predictions <- predict(nb_model, newdata = test_data)

# Create a confusion matrix
confusion_mat <- table(Actual = test_data$label, Predicted = predictions)
cat("Confusion Matrix:\n")
print(confusion_mat)

# Calculate accuracy
accuracy <- sum(diag(confusion_mat)) / sum(confusion_mat)
cat("\nModel Accuracy:", round(accuracy * 100, 2), "%\n")

5. Scaling to Full 250k Samples per Class

When you move to the full dataset, you might run into memory issues with a large DTM. Here are two fixes:

  • Remove sparse terms: Use removeSparseTerms() to drop words that appear in very few documents (e.g., keep terms that appear in at least 1% of documents):
    dtm_sparse <- removeSparseTerms(dtm, 0.99) # 0.99 = keep terms with sparsity ≤ 99%
    
  • Use a more efficient package: For large-scale text data, text2vec is faster and more memory-efficient than tm. It works seamlessly with Naive Bayes implementations too.

Quick Troubleshooting Tips

If your model performance is underwhelming:

  • Double-check your preprocessing: Randomly sample 5-10 preprocessed texts to confirm stopwords, punctuation, numbers, and irrelevant terms are removed, and stemming/lemmatization is applied correctly.
  • Adjust smoothing: If the model is overfitting (great on training, bad on test), increase the laplace parameter. If it's underfitting, decrease it.
  • Verify class balance: Even if your full dataset is balanced, ensure your train/test splits don't accidentally skew class ratios.

内容的提问来源于stack exchange,提问作者Peter Holler

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:01:23