如何调整R单因子预测模型实现双因子(办公室+月份)成交数预测
Great news—you’re already halfway there with the rpart model you’re using! Decision trees (like the one from rpart) are perfect for extending to multiple predictors, so adjusting the single-factor model to handle both Office and DealMonth is straightforward. Let’s walk through the fixes and improvements you need:
Step 1: Fix Column Name Mismatches
Looking at your code and dataset sample, there are a couple of column name inconsistencies that will cause errors:
- Your dataset uses
Officefor location, but your code references茂..DealOffice(likely a typo or encoding glitch from importing your CSV). - Your dataset has
CountDealsfor the target variable, but your code usesDealCount.
First, correct these references to match your actual CSV structure. Let’s assume your CSV columns are exactly as you listed: Office, DealMonth, DealYear, CountDeals.
Step 2: Ensure All Categorical Variables Are Factors
You already converted DealMonth to a factor, but Office should also be treated as a categorical variable (since it’s a list of locations, not numeric values). Add this line after loading your data:
dat$Office <- factor(dat$Office)
Step 3: Update the Model Formula to Include Both Predictors
The original article used a single predictor, but rpart lets you add multiple predictors with a simple formula change. Instead of a single variable, use both Office and DealMonth as predictors:
model <- rpart(CountDeals ~ Office + DealMonth, data = train_data)
This tells the model to learn patterns based on both office location and month—exactly what you need to make predictions like "Atlanta in January".
Step 4: Test and Predict with Your Target Input
Once the model is trained, you can predict for any combination of Office and DealMonth easily. For example, to get the predicted deal count for Atlanta in January (month 1):
predict(model, data.frame(Office = factor("Atlanta"), DealMonth = factor(1)))
Or for December (month 12) as in your sample code:
predict(model, data.frame(Office = factor("Atlanta"), DealMonth = factor(12)))
Quick Cleanup of Your Existing Code
Here’s how your corrected code should look, with all fixes applied:
dat <- read.csv("RawDataDealCountSummary.csv") head(dat) str(dat) # Convert categorical variables to factors dat$DealMonth <- factor(dat$DealMonth) dat$Office <- factor(dat$Office) # Split train/test data train_data <- dat[dat$DealYear < 2017,] test_data <- dat[dat$DealYear == 2017,] # Plot relationships (corrected column names) plot(dat$Office, dat$CountDeals) # Train model with both predictors model <- rpart(CountDeals ~ Office + DealMonth, data = train_data) # Test predictions test_counts <- test_data$CountDeals p <- predict(model, test_data) plot(p - test_counts) # Predict for specific office and month predict(model, data.frame(Office = factor('Atlanta'), DealMonth = factor(1)))
Why This Works
Decision trees (from rpart) natively handle categorical predictors like office locations and months. The model will learn patterns like "Atlanta averages X deals in January, Y deals in February" and "Dallas averages Z deals in January" from your training data, then use those patterns to make predictions for new combinations.
This approach builds directly on the single-factor model you learned—no need to switch to a completely different method!
内容的提问来源于stack exchange,提问作者walkery

