基于时间和聚类分组的多变量预测R代码需求
Hey there! As someone who’s worked with R for years and helped plenty of beginners, I’ll walk you through a straightforward, step-by-step solution to build your grouped prediction table and export it to CSV. Let’s keep this simple and actionable since you’re new to R.
First, we’ll use the tidyverse package—it’s a beginner-friendly toolkit that makes grouping data, running predictions, and exporting files way easier. Install and load it like this:
# Install the package if you haven't already install.packages("tidyverse") # Load it into your R session library(tidyverse)
Let’s make a test dataset that mirrors your example. This lets you test the code before using your real data:
set.seed(123) # Ensures random numbers are reproducible sample_data <- tibble( Time = rep(seq.Date(as.Date("2018-04-21"), as.Date("2018-04-25"), by = "day"), 2), Cluster = rep(c("A", "B"), each = 5), X1 = sample(10:80, 10, replace = TRUE), X2 = sample(30:90, 10, replace = TRUE), X3 = sample(20:80, 10, replace = TRUE) ) print(sample_data)
We’ll build a simple function that takes a single Cluster’s time-series data and generates predictions. For this example, we’ll use a lag-based prediction (predict the next day’s values using the most recent observed values)—you can swap this for a more complex model (like linear regression) later if needed:
# Function to predict the next time step for a single Cluster predict_group <- function(group_data) { # Get the last recorded date and calculate the next date last_time <- max(group_data$Time) next_time <- last_time + days(1) # Generate predictions using the latest observed values predictions <- group_data %>% slice_tail(n = 1) %>% # Grab the most recent row mutate( Time = next_time, X1_pred = X1, # Replace with model output if using regression X2_pred = X2, X3_pred = X3 ) %>% select(Time, Cluster, X1_pred, X2_pred, X3_pred) return(predictions) }
*Pro tip: If you want to use a regression model instead, here’s a quick snippet for X1:
model_x1 <- lm(X1 ~ Time, data = group_data) x1_pred <- predict(model_x1, newdata = tibble(Time = next_time)) ```* # Step 4: Apply Predictions to Every Cluster Now we’ll group the data by `Cluster`, run the prediction function on each group, and combine all results into one table: ```r # Group by Cluster, run predictions, and flatten the results prediction_results <- sample_data %>% group_by(Cluster) %>% nest() %>% # Nest each Cluster's data into a separate table mutate(predictions = map(data, predict_group)) %>% # Run predictions on each group unnest(predictions) %>% # Turn nested predictions into a flat table select(-data) # Remove the unused nested data column # Optional: Combine original data with predictions combined_data <- sample_data %>% bind_rows(prediction_results) %>% arrange(Time, Cluster) print(combined_data)
Finally, export your table (either just predictions or combined original + predictions) to a CSV file you can use later:
# Export combined data to CSV write_csv(combined_data, "prediction_table.csv") # Or export only the predictions: # write_csv(prediction_results, "only_predictions.csv")
- Predict multiple time steps: Add a loop inside
predict_group()to generate predictions for 2+ future days. - Use your real data: Replace
sample_datawith your actual dataset (load it withread_csv("your_real_data.csv")). - Add more variables: Extend the prediction logic to include X4, X5, etc.—just mirror the pattern for X1/X2/X3.
内容的提问来源于stack exchange,提问作者B. Alt

