You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于分年级线性模型生成预测值95%置信区间的技术问询

Solution: Generate Grade-Specific Prediction Intervals and Flag Outliers

Got it, let's walk through how to add 95% prediction intervals (for individual student scores) and use grade-specific standard deviations to flag outliers in your 2020 test data.

First, a quick clarification: when working with individual student predictions, prediction intervals are more appropriate than confidence intervals. Confidence intervals estimate the average score for a group, while prediction intervals account for both model uncertainty and individual student variation—perfect for your outlier-detection goal. I'll focus on prediction intervals here, but I'll note how to switch to confidence intervals if that's what you need.

Step 1: Update Your Model Object to Include Grade-Specific Standard Deviations

First, we'll modify your model code to extract the residual standard deviation (sigma) for each grade's model—this is your "grade-specific value" for setting outlier thresholds.

library(tidyverse)
library(broom)
library(purrr)

# Your existing data creation code (unchanged)
school_year <- rep(2017:2020, 120)
grade <- rep(1:12, each = 40)
attendance_rate <- round(runif(480, min=25, max=100), 1)
test_growth <- round(runif(480, min = -12, max = 38))
binary_flag <- round(runif(480, min = 0, max = 1))
score <- round(runif(480, min = 92, max = 370))
survey_response <- round(runif(480, min = 1, max = 4))
df <- data.frame(school_year, grade, attendance_rate, test_growth, binary_flag, score, survey_response)
df$survey_response[df$grade == 1] <- NA

df_train <- df %>% filter(!(school_year == 2020))
df_predict <- df %>% filter(school_year == 2020)

# Updated model code with grade-specific sigma
model <- df_train %>% 
  group_by(grade) %>% 
  nest() %>% 
  mutate(
    fit = map(data, ~ if(all(is.na(.x$survey_response))) 
                lm(score ~ attendance_rate + test_growth + binary_flag, data = .x) 
              else 
                lm(score ~ attendance_rate + test_growth + binary_flag + survey_response, data = .x)),
    tidied = map(fit, tidy),
    augmented = map(fit, augment),
    glanced = map(fit, glance),
    # Extract residual standard deviation for each grade's model
    grade_sigma = map_dbl(glanced, ~ .x$sigma)
  )

Step 2: Generate Predictions with 95% Intervals and Flag Outliers

Next, we'll generate predictions for the 2020 data, attach the prediction intervals, and use the grade-specific sigma to flag outliers (we'll use the common rule of thumb: a score is an outlier if it's more than 2 standard deviations away from its predicted value).

# Generate predictions with intervals and flag outliers
final_predictions <- df_predict %>% 
  nest(test_data = -grade) %>% 
  inner_join(model, by = 'grade') %>% 
  mutate(
    # Generate predictions with 95% prediction intervals (default level = 0.95)
    pred_output = map2(fit, test_data, ~ predict(.x, newdata = .y, interval = "prediction")),
    # Convert the prediction matrix to a tidy data frame
    pred_df = map(pred_output, ~ as.data.frame(.x) %>% 
                    rename(pred_score = fit, lwr_pred = lwr, upr_pred = upr)),
    # Combine test data, predictions, and grade-specific sigma
    combined_results = pmap(list(test_data, pred_df, grade_sigma), ~ bind_cols(..1, ..2, grade_sigma = ..3))
  ) %>% 
  # Unnest to get one row per student
  unnest(combined_results) %>% 
  # Calculate score difference and flag outliers
  mutate(
    score_diff = score - pred_score,
    # Mark as outlier if absolute difference exceeds 2*grade-specific sigma
    is_outlier = abs(score_diff) > 2 * grade_sigma
  )

Key Details to Note:

  • Prediction vs. Confidence Intervals: If you actually need confidence intervals (for the average score of students with those attributes), replace interval = "prediction" with interval = "confidence". Prediction intervals are wider because they account for individual student variation, which is better for identifying outliers.
  • Outlier Threshold: We used 2 * grade_sigma as the threshold, but you can adjust this (e.g., 1.96 for exact 95% coverage, or 3 for stricter outlier detection).
  • Handling NA Values: Your code already accounts for missing survey_response in 1st grade, and the prediction step will automatically use the correct model (without survey_response) for 1st grade students in the test set.

Example Output

The final_predictions data frame will have all your original test data columns, plus:

  • pred_score: The predicted score for the student
  • lwr_pred/upr_pred: 95% prediction interval bounds
  • grade_sigma: Residual standard deviation for the student's grade model
  • score_diff: Actual score minus predicted score
  • is_outlier: Boolean flag indicating if the student's score is an outlier

内容的提问来源于stack exchange,提问作者ra_learns

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 14:12:35