基于R语言随机森林模型的事件工单优先级分类实践问询
Hey there! I see you're just getting started with R and machine learning, working on predicting incident ticket priorities from their descriptions—great project! Let's build out your workflow step by step, starting from where you left off.
Your Current Progress
You've already loaded the key libraries and your dataset, which is a perfect starting point:
library(tm) library(SnowballC) library(caTools) library(randomForest) incidents = read.csv("incident.csv", stringsAsFactors = FALSE) str(incidents)
Partial output you shared:
'data.frame': 4265 obs. of 7 variables:
$ number : chr "INC0031193" "INC0037867" "INC0159979" "INC0031446" ...
[remaining variables omitted]
Now let's walk through the critical next steps to turn that raw text data into a working predictive model.
Step 1: Text Preprocessing (Make Text Machine-Friendly)
Since your predictor is ticket description text, we need to clean and structure it into numerical features that the random forest can use. Here's how to do that:
# Replace "description" with the actual name of your text column in the dataset corpus = VCorpus(VectorSource(incidents$description)) # Convert all text to lowercase corpus = tm_map(corpus, content_transformer(tolower)) # Remove punctuation marks corpus = tm_map(corpus, removePunctuation) # Strip out common stopwords (e.g., "the", "and") that don't add predictive value corpus = tm_map(corpus, removeWords, stopwords("english")) # Stem words (reduce to root form, e.g., "running" → "run") corpus = tm_map(corpus, stemDocument) # Convert cleaned text to a Document-Term Matrix (text → numerical features) dtm = DocumentTermMatrix(corpus) # Remove rare words to avoid overfitting (keeps terms appearing in at least 1% of tickets) dtm_sparse = removeSparseTerms(dtm, 0.99) # Convert to a data frame compatible with modeling text_features = as.data.frame(as.matrix(dtm_sparse)) # Add your target variable (priority) back to the feature set text_features$priority = incidents$priority
Step 2: Prep Data for Training
Random Forest requires categorical target variables to be factors, so let's fix that and split our data into training/test sets:
# Convert priority to a factor (critical for classification) text_features$priority = as.factor(text_features$priority) # Set seed for reproducible results set.seed(123) # Split data: 70% training, 30% testing split = sample.split(text_features$priority, SplitRatio = 0.7) train_data = subset(text_features, split == TRUE) test_data = subset(text_features, split == FALSE)
Step 3: Train the Random Forest Model
Now we can build and train our classifier:
set.seed(456) # ntree = number of trees in the forest (adjust based on performance) rf_model = randomForest(priority ~ ., data = train_data, ntree = 200) # Check model summary to see initial performance print(rf_model)
Step 4: Evaluate Model Performance
Let's test how well our model works on unseen data:
# Generate predictions on the test set predictions = predict(rf_model, newdata = test_data) # Create a confusion matrix to see class-wise performance confusion_matrix = table(test_data$priority, predictions) print(confusion_matrix) # Calculate overall accuracy accuracy = sum(diag(confusion_matrix)) / sum(confusion_matrix) cat("Model Accuracy:", round(accuracy * 100, 2), "%\n")
Quick Tips to Boost Performance
- Tune Hyperparameters: Use
tuneRF()to find the optimalmtry(number of variables sampled per tree split) for better results. - Fix Imbalanced Data: If some priority classes are underrepresented, try oversampling minority classes or using the
classwtparameter inrandomForest()to assign weights. - Add More Features: Include other relevant columns (like ticket category, submitter department) alongside text features to give the model more context.
内容的提问来源于stack exchange,提问作者Sourav94

