基于分年级线性模型生成预测值95%置信区间的技术问询
Got it, let's walk through how to add 95% prediction intervals (for individual student scores) and use grade-specific standard deviations to flag outliers in your 2020 test data.
First, a quick clarification: when working with individual student predictions, prediction intervals are more appropriate than confidence intervals. Confidence intervals estimate the average score for a group, while prediction intervals account for both model uncertainty and individual student variation—perfect for your outlier-detection goal. I'll focus on prediction intervals here, but I'll note how to switch to confidence intervals if that's what you need.
Step 1: Update Your Model Object to Include Grade-Specific Standard Deviations
First, we'll modify your model code to extract the residual standard deviation (sigma) for each grade's model—this is your "grade-specific value" for setting outlier thresholds.
library(tidyverse) library(broom) library(purrr) # Your existing data creation code (unchanged) school_year <- rep(2017:2020, 120) grade <- rep(1:12, each = 40) attendance_rate <- round(runif(480, min=25, max=100), 1) test_growth <- round(runif(480, min = -12, max = 38)) binary_flag <- round(runif(480, min = 0, max = 1)) score <- round(runif(480, min = 92, max = 370)) survey_response <- round(runif(480, min = 1, max = 4)) df <- data.frame(school_year, grade, attendance_rate, test_growth, binary_flag, score, survey_response) df$survey_response[df$grade == 1] <- NA df_train <- df %>% filter(!(school_year == 2020)) df_predict <- df %>% filter(school_year == 2020) # Updated model code with grade-specific sigma model <- df_train %>% group_by(grade) %>% nest() %>% mutate( fit = map(data, ~ if(all(is.na(.x$survey_response))) lm(score ~ attendance_rate + test_growth + binary_flag, data = .x) else lm(score ~ attendance_rate + test_growth + binary_flag + survey_response, data = .x)), tidied = map(fit, tidy), augmented = map(fit, augment), glanced = map(fit, glance), # Extract residual standard deviation for each grade's model grade_sigma = map_dbl(glanced, ~ .x$sigma) )
Step 2: Generate Predictions with 95% Intervals and Flag Outliers
Next, we'll generate predictions for the 2020 data, attach the prediction intervals, and use the grade-specific sigma to flag outliers (we'll use the common rule of thumb: a score is an outlier if it's more than 2 standard deviations away from its predicted value).
# Generate predictions with intervals and flag outliers final_predictions <- df_predict %>% nest(test_data = -grade) %>% inner_join(model, by = 'grade') %>% mutate( # Generate predictions with 95% prediction intervals (default level = 0.95) pred_output = map2(fit, test_data, ~ predict(.x, newdata = .y, interval = "prediction")), # Convert the prediction matrix to a tidy data frame pred_df = map(pred_output, ~ as.data.frame(.x) %>% rename(pred_score = fit, lwr_pred = lwr, upr_pred = upr)), # Combine test data, predictions, and grade-specific sigma combined_results = pmap(list(test_data, pred_df, grade_sigma), ~ bind_cols(..1, ..2, grade_sigma = ..3)) ) %>% # Unnest to get one row per student unnest(combined_results) %>% # Calculate score difference and flag outliers mutate( score_diff = score - pred_score, # Mark as outlier if absolute difference exceeds 2*grade-specific sigma is_outlier = abs(score_diff) > 2 * grade_sigma )
Key Details to Note:
- Prediction vs. Confidence Intervals: If you actually need confidence intervals (for the average score of students with those attributes), replace
interval = "prediction"withinterval = "confidence". Prediction intervals are wider because they account for individual student variation, which is better for identifying outliers. - Outlier Threshold: We used
2 * grade_sigmaas the threshold, but you can adjust this (e.g., 1.96 for exact 95% coverage, or 3 for stricter outlier detection). - Handling NA Values: Your code already accounts for missing
survey_responsein 1st grade, and the prediction step will automatically use the correct model (withoutsurvey_response) for 1st grade students in the test set.
Example Output
The final_predictions data frame will have all your original test data columns, plus:
pred_score: The predicted score for the studentlwr_pred/upr_pred: 95% prediction interval boundsgrade_sigma: Residual standard deviation for the student's grade modelscore_diff: Actual score minus predicted scoreis_outlier: Boolean flag indicating if the student's score is an outlier
内容的提问来源于stack exchange,提问作者ra_learns

